Omni-Decision is an omni-modal agent that uses evidence-ledger planning to manage noisy multimodal observations. It replaces the growing dialogue history with a compact evidence ledger that tracks missing, confirmed, and conflicting evidence, allowing the planner to operate on a streamlined context. The system achieves state‑of‑the‑art accuracy on OmniGAIA and WorldSense while operating at a fraction of the cost of larger models.
By Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Yuhao Wang, Junhan Shi, Lingrui Mei, Tianming Yang, Steven Hoi
arXiv:2608.31005v1 Announce Type: new
Abstract: Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurr...
By Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Ruirui Li
arXiv:2609.15128v1 Announce Type: new
Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...
By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou
The paper introduces PACE, a factor-guided, progressive framework for acquiring evidence in long-video question answering. PACE first indexes clip-level descriptions using question-derived factors, then refines evidence retrieval with contrastive cues derived from candidate answers. On the MMR‑V dataset, PACE achieves 42.6% accuracy and recovers 66.9% of annotated cues, outperforming direct inference and prior agentic baselines, and shows consistent improvements across several long-video benchmarks.
By Baixuan Xu, Yinyui Xu, Tianshi Zheng, Zhaowei Wang, Weiqi Wang, Haochen Shi, Jiayu Liu, Qing Zong, Xiyu Ren, Xinyu Geng, Zhitao He, Yangqiu Song
arXiv:2602. 22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world.
By Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities inde...
arXiv:2608.21868v1 Announce Type: new
Abstract: Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile....
By Ao Chen, Xiaojiang Peng
arXiv:2608.05592v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budge...
By Ziling Huang, Shin'ichi Satoh
arXiv:2609.14823v1 Announce Type: new
Abstract: Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiol...
By Ji Lu, Lifei Liu, Haoran Yu, Xianglong Wang, Yiru Fang, Kuo Yang, Huiran Duan, Jianping Gou
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information.
arXiv:2609.39566v1 Announce Type: new
Abstract: Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected e...
By Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
arXiv:2607. 19793v1 Announce Type: new Abstract: Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions.
By Zhengxian Wu, Junjie Gao, Kai Yang