Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
-
Updated
Aug 21, 2024 - Python
Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
[AAAI 2022] Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
[ICLR 2025] TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
Pytorch implementation of the paper 'Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding' (AAAI2024).
Official Implementation of Moment Alignment Transformer
[NCA] Official implementation of the paper Motion2Language, Unsupervised learning of synchronized semantic motion segmentation
[BMVC 2024] Official Implementation of the paper guided attention for interpretable motion captioning
Transformer with Controlled Attention for Synchronous Motion Captioning
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Offline audio QA, transcription, translation, emotion, temporal grounding and timed summaries on Apple Silicon — native GigaChat Audio MLX.
Corpus-scale video moment retrieval benchmark — 75h of rights-cleared video, 500 queries, 4 difficulty tiers. Built on Mixpeek.
EMCompress: Video-LLMs with Endomorphic Multimodal Compression (ACL 2026 Findings)
Semantic Augmentation and Gated-adapter Encoding for Video Moment Retrieval — closing the linguistic robustness gap with LLM-driven query augmentation and nested gated-adapters.
Multimodal video analysis — Whisper transcription, LLM chapter summarisation, CLIP key-frame extraction, and natural language Q&A over uploaded video content.
Temporal context grounding for LLM sessions. Stops AI from living in the past.
SPIRAL: structured supervision harvesting and self-refining inference for surgical video understanding. ECCV 2026 MedVidU (Oral).
To associate your repository with the temporal-grounding topic, visit your repo's landing page and select "manage topics."