arXiv Computation and Language By Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji

Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models

Read the original on arXiv Computation and Language →

The paper introduces Reward‑Tilted On‑Policy Distillation (RT‑OPD), a method that enhances acoustic grounding in audio‑language models by using a frozen teacher to generate a reward based on the contrast between token predictions with and without audio. This reward reshapes the teacher distribution for reverse‑KL distillation, encouraging students to rely more on acoustic evidence. Experiments on two compact students across three benchmarks show that RT‑OPD consistently outperforms vanilla OPD, and a 3B model trained with RT‑OPD achieves 72.72% accuracy on the MMAU benchmark, surpassing other 3B models and rivaling larger 7B and 8B models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 24

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

arXiv:2607. 21550v1 Announce Type: new Abstract: While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data.

By Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin
arXiv AI
Sep 7

Tracing Audio Grounding and Answer Selection in Audio LLMs

The paper investigates how Audio Large Language Models (Audio LLMs) actually use audio input to determine answers, rather than relying on textual cues. It finds that replacing audio with silence or unrelated audio degrades performance more after training than before, that acoustic information shapes representations in early-to-middle layers and influences final predictions in middle-to-late layers, and that training impacts specific layer bands most strongly. These observations offer a mechanistic view of how training enhances the use of acoustic evidence in Audio LLMs.

By Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung
arXiv Machine Learning
1d ago

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.

By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
arXiv Computer Vision
3d ago

OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

OP-CAD introduces a curriculum-based, on-policy clean-audio distillation framework that enhances audio-visual reasoning under environmental noise and competing speech. The method trains a student model from mild to severe noise, using a frozen teacher that provides token-level supervision based on clean audio and verified answers, while selectively weighting positions sensitive to acoustic interference. Experiments show OP‑CAD outperforms existing methods across all noise conditions, preserving clean‑correct answers without sacrificing overall accuracy.

By Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang