VideoResearcher is a training‑free, multi‑agent framework that autonomously designs, tests, and refines high‑impact tools for long‑video understanding. It operates through dual Solving and Evolving loops, analyzing tool‑use trajectories to identify gaps, coordinating specialized agents to develop and validate executable tools, and reusing evolved tools to improve evidence acquisition in subsequent reasoning. The approach achieves state‑of‑the‑art performance among self‑improving agents and approaches the human‑designed upper bound, demonstrating a paradigm that expands agent capabilities while reducing costly manual engineering.
By Dingqiang Ye, Dongdi Zhao, Kaishen Wang, Qingqiao Hu, Jingchen Sun, Yijun Liang, Yuqi Jia, Yiqiao Huang, Yunjie Tian, Jiaxing Zhang, Chuanyang Jin, Ke Zhang, Vishal M. Patel, Di Fu
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched.
VideoHarness‑RSI explores how improving the executable context‑construction program alone can enhance long‑video understanding with frozen vision‑language models. By recursively searching for better harnesses—programs that select and structure video segments—using an outer‑loop proposer that learns from prior programs and execution traces, the method consistently outperforms weaker hand‑crafted baselines and further improves upon stronger ones. The resulting harnesses transfer to other long‑video benchmarks without additional search, demonstrating that executable context construction is a distinct, reusable optimization layer.
By Guoyang Xu, Hao Chen
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
By Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
arXiv:2609.09985v1 Announce Type: new
Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a...
By Sheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao
arXiv:2609.12818v1 Announce Type: new
Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal...
By Sen Yang, Boqiang Duan, Jing Yang, Weihao Bo, Jie Liu, Boyuan Tong, Ze Feng, Wenkang Zhang, Jingdong Wang, Hua Wu
arXiv:2606. 26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable.
By Yangfan He, Yujin Choi, Jaehong Yoon
arXiv:2609.15478v1 Announce Type: new
Abstract: Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmar...
By Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi, Pinxin Liu, Yicheng Wang, Yunzhong Xiao, Zhangyun Tan, Zeliang Zhang, Chao Huang, Susan Liang, Qianxiang Shen, Luchuan Song, Ali Vosoughi, Mingqian Feng, Melika Filvantorkaman, Chenliang Xu
arXiv:2605. 21917v2 Announce Type: replace-cross Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support.
By Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng, Vidya Murali
arXiv:2603. 26266v3 Announce Type: replace Abstract: Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction.
By Rui Xie, Zhi Gao, Chenrui Shi, Zirui Shang, Lu Chen, Qing Li
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.
By Yanxiang Huang, Guohua Gao, Zhaoyang Wei
arXiv:2606. 27922v1 Announce Type: cross Abstract: Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters.
By Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, Fei Ma