Video-RAG Agent
Ask natural-language questions about your videos — Video-RAG searches spoken dialogue, visuals, on-screen text, and motion across your library — chaining tools together when a question needs more than one — to find and verify the answer.
New here? Press play
Play this quick 1-minute walkthrough to get started instantly.
Configuration
Agent LLM
VLM (visual verification)
If this model fails mid-query, the agent automatically tries the next available VLM and logs the attempt in the Tool Trace tab.
Session
Chat with the Video-RAG Agent
Tool Call Trace
Shows the tool calls the agent made for each of your messages (newest turn on top), including tool input/output and any VLM fallback attempts if a provider failed mid-call.
No tool calls yet this session.
Video Library
Pre-loaded sample videos
Your uploaded videos (this session only)
Upload one or more videos to build your session library.
Model Information
Every model this project relies on, mapped to the tool it powers. Accuracy figures are published results from each model's own paper, model card, or an independent community evaluation — none of these were re-measured against this project's own video library.
Per-tool local models (run directly on this Space)
Agent & reasoning models (called via API)
Note: this project runs entirely on free-tier infrastructure — Hugging Face ZeroGPU for the local models above, and free-tier inference APIs for the agent and vision-language models. Speed, availability, and rate limits are all subject to those free-tier constraints rather than the models' own ceiling.