Multimodal-VideoRAG: Using BridgeTower Embeddings and Large Vision Language Models
-
Updated
Sep 26, 2024 - Jupyter Notebook
Multimodal-VideoRAG: Using BridgeTower Embeddings and Large Vision Language Models
Multimodal RAG over video: fine-tuned BridgeTower embeddings + LanceDB retrieval + LLaVA-1.5 grounded answers, with a Gradio UI that returns the matching video clip
To associate your repository with the bridgetower topic, visit your repo's landing page and select "manage topics."