arXiv AI

Measuring the Redundancy of Decoder Layers in SpeechLLMs

arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.

arXiv AI
Sep 12

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

X-AuT is a progressive framework for compressing the audio encoder of speech large language models. It selects layer combinations via short behavioral probes and restores performance through representation alignment, cross‑scale distillation, scheduled student‑policy supervision, and LoRA finetuning, while keeping the language‑model backbone frozen. On ten Chinese–English benchmarks, reducing Qwen3‑ASR‑0.6B’s encoder from 18 to 16 layers lowers macro‑average error from 5.61% to 5.27%, and a 14‑layer model achieves 5.75% error with 20.7% fewer parameters.

By Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu Huang
arXiv AI
Sep 18

Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.

By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
arXiv Machine Learning
Sep 24

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."

By Rasmus Aagaard, Nicki Skafte Detlefsen
arXiv Computation and Language
Sep 22

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

arXiv:2604.06487v2 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...

By Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv AI
Jun 10

Whisfusion: Parallel ASR Decoding with Masked Diffusion

arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.

By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim