Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities inde...
OmniReasoning introduces a new benchmark, OmniReasoningBench, that requires both audio and visual evidence for answering 1,150 multiple-choice and open-ended questions across two tasks. The authors also develop OmniQA, a data engine that automatically generates evidence‑grounded QA pairs with time‑stamped clue chains, producing training datasets OmniReasoning‑SFT‑112K and OmniReasoning‑RL‑19K. Finally, they propose Modality‑Factored Self‑Distillation (MFSD), an on‑policy self‑distillation method that assigns token‑level credit by evaluating responses under modality‑specific clue contexts, enabling the OmniReasoning‑30B‑A3B model to achieve significant performance gains on both the new benchmark and existing video benchmarks.
By Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
arXiv:2610.02181v1 Announce Type: new
Abstract: We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with nativ...
By Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu
arXiv:2609.23589v1 Announce Type: cross
Abstract: Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio repr...
By Jiaheng Dong, Xiaofeng Yu, Jean Honorio, Abhirup Ghosh, Hong Jia, Ting Dang
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.
By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv:2607. 14682v1 Announce Type: new Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge.
By Harikrishnan P M, Goutham Vignesh, Ganesh Parab, Saisubramaniam Gopalakrishnan, Vishal Vaddina, Varun V, Rohit Agrawal
arXiv:2603.02266v2 Announce Type: replace-cross
Abstract: Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-La...
By Ruixiang Mao, Xiangnan Ma, Dan Chen, Ziming Zhu, Yuan Ge, Aokai Hao, Haishu Zhao, Yifu Huo, Qing Yang, Kaiyan Chang, Xiaoqian Liu, Chenglong Wang, Qiaozhi He, Tong Xiao, Jingbo Zhu
The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning.
"whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."
By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
By Jingpei Wu, Xiao Han, Weixiang Shen, Boer Zhang, Zifeng Ding, Volker Tresp
arXiv:2608. 03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction.
By Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang
OmniVChat defines a native audio‑visual dialogue task where models receive raw audio and video from a user and produce text responses, eliminating the need for separate text queries or speech recognition. To address data scarcity and evaluation challenges, the authors introduce OmniVChat‑Studio, a multi‑agent engine that synthesizes single‑ and multi‑turn dialogues, and OmniVChat‑Bench, a benchmark assessing models across five dialogue ability categories. They also propose OmniVChat‑RL, a reinforcement‑learning reward that balances reply correctness, efficiency, and style, and demonstrate that training Qwen3‑Omni‑Instruct with this reward on synthesized data improves performance on both synthetic and human‑recorded benchmarks.
By Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu, Zining Liang, Zhengxi Liu, Xiquan Li, Xie Chen, Xize Cheng, Qize Yang, Jin Xu, Qiuqiang Kong
The paper investigates why post‑training boosts reasoning more than perception in vision‑language models. Using a diagnostic framework with synthetic tasks, it finds a perception‑reasoning asymmetry: supervised fine‑tuning suffers from token imbalance, while reinforcement learning suffers from reward coupling. The authors propose reweighting losses and perception‑aware rewards, achieving up to 18.2‑point and 6.0‑point gains respectively, and show that these methods also improve real‑world visual reasoning by up to 3.3 points.
By Xueqing Wu, Yu-Chi Lin, Kai-Wei Chang, Nanyun Peng