EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
X-AuT is a progressive framework for compressing the audio encoder of speech large language models. It selects layer combinations via short behavioral probes and restores performance through representation alignment, cross‑scale distillation, scheduled student‑policy supervision, and LoRA finetuning, while keeping the language‑model backbone frozen. On ten Chinese–English benchmarks, reducing Qwen3‑ASR‑0.6B’s encoder from 18 to 16 layers lowers macro‑average error from 5.61% to 5.27%, and a 14‑layer model achieves 5.75% error with 20.7% fewer parameters.
OP-CAD introduces a curriculum-based, on-policy clean-audio distillation framework that enhances audio-visual reasoning under environmental noise and competing speech. The method trains a student model from mild to severe noise, using a frozen teacher that provides token-level supervision based on clean audio and verified answers, while selectively weighting positions sensitive to acoustic interference. Experiments show OP‑CAD outperforms existing methods across all noise conditions, preserving clean‑correct answers without sacrificing overall accuracy.
The paper introduces Reward‑Tilted On‑Policy Distillation (RT‑OPD), a method that enhances acoustic grounding in audio‑language models by using a frozen teacher to generate a reward based on the contrast between token predictions with and without audio. This reward reshapes the teacher distribution for reverse‑KL distillation, encouraging students to rely more on acoustic evidence. Experiments on two compact students across three benchmarks show that RT‑OPD consistently outperforms vanilla OPD, and a 3B model trained with RT‑OPD achieves 72.72% accuracy on the MMAU benchmark, surpassing other 3B models and rivaling larger 7B and 8B models.