arXiv Machine Learning

Analyzing the Importance of Blank for CTC-Based Knowledge Distillation

arXiv:2506. 01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear.

arXiv Machine Learning
4d ago

Data Unlearning via Inverse Distillation

arXiv:2609.36099v1 Announce Type: new Abstract: Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce u...

By Aleksei Leonov, Nikita Kornilov, Zhenhe Zhang, Evgeny Burnaev, Iaroslav Koshelev, Alexander Korotin
arXiv Computation and Language
Sep 25

Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models

The paper introduces Reward‑Tilted On‑Policy Distillation (RT‑OPD), a method that enhances acoustic grounding in audio‑language models by using a frozen teacher to generate a reward based on the contrast between token predictions with and without audio. This reward reshapes the teacher distribution for reverse‑KL distillation, encouraging students to rely more on acoustic evidence. Experiments on two compact students across three benchmarks show that RT‑OPD consistently outperforms vanilla OPD, and a 3B model trained with RT‑OPD achieves 72.72% accuracy on the MMAU benchmark, surpassing other 3B models and rivaling larger 7B and 8B models.

By Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji
Hugging Face Trending Papers
Jun 3

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.

arXiv AI
Jul 7

Unified Audio Intelligence Without Regressing on Text Intelligence

arXiv:2607. 05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech.

By Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping