Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
-
Updated
Aug 21, 2024 - Python
Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
[AAAI 2022] Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
[ICLR 2025] TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
paper list on Video Moment Retrieval (VMR), or Natural Language Video Localization (NLVL), or Temporal Sentence Grounding in Videos (TSGV))
Pytorch implementation of the paper 'Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding' (AAAI2024).
Official Implementation of Moment Alignment Transformer
[NCA] Official implementation of the paper Motion2Language, Unsupervised learning of synchronized semantic motion segmentation
[BMVC 2024] Official Implementation of the paper guided attention for interpretable motion captioning
Transformer with Controlled Attention for Synchronous Motion Captioning
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Offline audio QA, transcription, translation, emotion, temporal grounding and timed summaries on Apple Silicon — native GigaChat Audio MLX.
Corpus-scale video moment retrieval benchmark — 75h of rights-cleared video, 500 queries, 4 difficulty tiers. Built on Mixpeek.
EMCompress: Video-LLMs with Endomorphic Multimodal Compression (ACL 2026 Findings)
Semantic Augmentation and Gated-adapter Encoding for Video Moment Retrieval — closing the linguistic robustness gap with LLM-driven query augmentation and nested gated-adapters.
Multimodal video analysis — Whisper transcription, LLM chapter summarisation, CLIP key-frame extraction, and natural language Q&A over uploaded video content.
Temporal context grounding for LLM sessions. Stops AI from living in the past.
This paper presents the VLMI framework to detect activities in complex videos. It combines Swin Transformer video features with language prompts and an EIoU-based similarity measure, enabling accurate, query-driven activity detection and timestamping, handling visual noise and temporal uncertainty without full manual labeling.
Structured Qwen3-VL video timeline extraction and temporal grounding evaluation runtime.
To associate your repository with the temporal-grounding topic, visit your repo's landing page and select "manage topics."