Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the...
X-AuT is a progressive framework for compressing the audio encoder of speech large language models. It selects layer combinations via short behavioral probes and restores performance through representation alignment, cross‑scale distillation, scheduled student‑policy supervision, and LoRA finetuning, while keeping the language‑model backbone frozen. On ten Chinese–English benchmarks, reducing Qwen3‑ASR‑0.6B’s encoder from 18 to 16 layers lowers macro‑average error from 5.61% to 5.27%, and a 14‑layer model achieves 5.75% error with 20.7% fewer parameters.
By Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu Huang
arXiv:2511.08092v2 Announce Type: replace-cross
Abstract: We challenge the conventional view of neural network pruning as solely a compression technique, demonstrating that one-shot magnitude pruning...
By Julian Irigoyen, Arthur S\"ohler, Andreas S{\o}eborg Kirkedal
arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.
By Adel Moumen, Guangzhi Sun, Philip C Woodland
arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.
By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim
arXiv:2509.10452v3 Announce Type: replace-cross
Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance....
By Akshat Pandey, Karun Kumar, Raphael Tang