CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
VideoResearcher is a training‑free, multi‑agent framework that autonomously designs, tests, and refines high‑impact tools for long‑video understanding. It operates through dual Solving and Evolving loops, analyzing tool‑use trajectories to identify gaps, coordinating specialized agents to develop and validate executable tools, and reusing evolved tools to improve evidence acquisition in subsequent reasoning. The approach achieves state‑of‑the‑art performance among self‑improving agents and approaches the human‑designed upper bound, demonstrating a paradigm that expands agent capabilities while reducing costly manual engineering.
arXiv:2609.12818v1 Announce Type: new Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal...
arXiv:2606. 27922v1 Announce Type: cross Abstract: Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters.
arXiv:2608. 07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments.
VisionCoach is an input‑adaptive reinforcement learning framework that enhances spatio‑temporal grounding in video reasoning by using visual prompting during training. The system selectively applies visual prompts to challenging inputs, amplifying question‑relevant evidence and suppressing distractors, and then internalizes these improvements through self‑distillation so that inference can be performed on raw videos without prompts. Experiments on multiple benchmarks (V‑STAR, VideoMME, World‑Sense, VideoMMMU, PerceptionTest, and Charades‑STA) show that VisionCoach achieves state‑of‑the‑art performance while maintaining a single efficient inference pathway.
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.