Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
VideoResearcher is a training‑free, multi‑agent framework that autonomously designs, tests, and refines high‑impact tools for long‑video understanding. It operates through dual Solving and Evolving loops, analyzing tool‑use trajectories to identify gaps, coordinating specialized agents to develop and validate executable tools, and reusing evolved tools to improve evidence acquisition in subsequent reasoning. The approach achieves state‑of‑the‑art performance among self‑improving agents and approaches the human‑designed upper bound, demonstrating a paradigm that expands agent capabilities while reducing costly manual engineering.
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched.
VideoHarness‑RSI explores how improving the executable context‑construction program alone can enhance long‑video understanding with frozen vision‑language models. By recursively searching for better harnesses—programs that select and structure video segments—using an outer‑loop proposer that learns from prior programs and execution traces, the method consistently outperforms weaker hand‑crafted baselines and further improves upon stronger ones. The resulting harnesses transfer to other long‑video benchmarks without additional search, demonstrating that executable context construction is a distinct, reusable optimization layer.
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
arXiv:2609.09985v1 Announce Type: new Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a...
arXiv:2609.12818v1 Announce Type: new Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal...