The paper investigates how post‑training compression techniques—such as pruning, quantization, and distillation—affect demographic fairness in Whisper speech‑recognition models. It finds that pruning and INT4 quantization significantly widen word‑error‑rate gaps between demographic groups, especially for Black/AA and Asian speakers, while distillation tends to reduce these gaps. The study introduces a temporal‑taxation metric to quantify the increased correction effort required for marginalized speakers after compression.
By Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers, Srinivasan Parthasarathy
arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that...
Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the...
The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%.
"whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."
By Rasmus Aagaard, Nicki Skafte Detlefsen
arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.
By Adel Moumen, Guangzhi Sun, Philip C Woodland
The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.
By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
arXiv:2606. 11836v1 Announce Type: cross Abstract: This paper presents a novel data-free and training-free compression approach for speech foundation models using channelwise clustering via k-means.
By Haoning Xu, Zhaoqing Li, Huimeng Wang, Youjun Chen, Chengxi Deng, Mengzhe Geng, Xunying Liu
arXiv:2609.27382v1 Announce Type: cross
Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents...
By Kian Shamsaie, Iman Modarressi
The paper investigates how to close the quality gap in low‑resource text‑to‑speech for Khmer and Korean using the VoxCPM2 model. By training a single low‑rank adaptation (LoRA) adapter on a shared 25.5‑hour corpus, the authors improve Khmer’s mean opinion score from 3.85 to 4.23 with a rank‑64 adapter, while Korean shows no significant gain. The study highlights that adaptation benefits mainly when the base model is weak and that training loss does not always align with human ratings.
By Phannet Pov, Hyun Woo Park, Voneat Pen, Sovandara Chhoun, Wan-Sup Cho, Saksonita Khoeurn
arXiv:2608.30927v1 Announce Type: cross
Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language...
By Chanhee Cho, Junhyuk Choi, Bugeun Kim
X-AuT is a progressive framework for compressing the audio encoder of speech large language models. It selects layer combinations via short behavioral probes and restores performance through representation alignment, cross‑scale distillation, scheduled student‑policy supervision, and LoRA finetuning, while keeping the language‑model backbone frozen. On ten Chinese–English benchmarks, reducing Qwen3‑ASR‑0.6B’s encoder from 18 to 16 layers lowers macro‑average error from 5.61% to 5.27%, and a 14‑layer model achieves 5.75% error with 20.7% fewer parameters.
By Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu Huang