arXiv Computer Vision

OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

OP-CAD introduces a curriculum-based, on-policy clean-audio distillation framework that enhances audio-visual reasoning under environmental noise and competing speech. The method trains a student model from mild to severe noise, using a frozen teacher that provides token-level supervision based on clean audio and verified answers, while selectively weighting positions sensitive to acoustic interference. Experiments show OP‑CAD outperforms existing methods across all noise conditions, preserving clean‑correct answers without sacrificing overall accuracy.

arXiv AI
Jun 2

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

arXiv:2606. 00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raising concerns about privacy and broader safety risks.

By Junkai Chen, Yuhao He, Junxiang You, Ruiqi Liu, Chenyu Wang, Shu Wu
arXiv Machine Learning
Jul 24

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

arXiv:2607. 21550v1 Announce Type: new Abstract: While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data.

By Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin
arXiv AI
3d ago

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a new benchmark, OmniReasoningBench, that requires both audio and visual evidence for answering 1,150 multiple-choice and open-ended questions across two tasks. The authors also develop OmniQA, a data engine that automatically generates evidence‑grounded QA pairs with time‑stamped clue chains, producing training datasets OmniReasoning‑SFT‑112K and OmniReasoning‑RL‑19K. Finally, they propose Modality‑Factored Self‑Distillation (MFSD), an on‑policy self‑distillation method that assigns token‑level credit by evaluating responses under modality‑specific clue contexts, enabling the OmniReasoning‑30B‑A3B model to achieve significant performance gains on both the new benchmark and existing video benchmarks.

By Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
arXiv AI
Sep 16

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.

By Chenhao Qiu, Dawei Li, Yechao Zhang, Lei Gong, Zhen Tan
arXiv Computer Vision
Sep 7

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.

By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv Computer Vision
Aug 28

Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models

Video-OPSD introduces a post‑training framework for Video Large Language Models that leverages privileged visual evidence to enhance on‑policy self‑distillation. The method constructs a self‑teacher conditioned only on annotated evidence frames, while the student processes the full video, allowing the teacher to provide more focused supervision. Additionally, an evidence‑guided token optimization scheme weights distillation based on each token’s reliance on privileged evidence, improving perceptually grounded reasoning. Experiments demonstrate consistent gains over standard OPSD and comparable performance to GRPO with less training time.

By Ziyue Wang, Shiqi Huang, Weiwen Xu, Bihan Wen, Xudong Jiang