The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.
By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
arXiv:2609.17981v1 Announce Type: cross
Abstract: Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-R...
By Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer Alwan
arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.
By Adel Moumen, Guangzhi Sun, Philip C Woodland
arXiv:2606. 11766v1 Announce Type: cross Abstract: Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments.
By Eungbeom Kim, Kyogu Lee
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
arXiv:2608.27783v3 Announce Type: replace-cross
Abstract: Speech language models (speech LLMs) can generate plausible outputs from audio that contains no usable speech evidence. We study this failure...
By Mengzhe Geng
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
arXiv:2606. 10233v1 Announce Type: cross Abstract: While speech quality is typically assessed on complete utterances, streaming and generative systems require incremental estimation from partial audio.
By Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, Shinji Watanabe
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
The paper introduces a reparameterization technique that injects feature noise to jointly optimize speech model performance and computational complexity during training. Unlike traditional pruning, this method dynamically adjusts model size for a desired performance‑complexity trade‑off without heuristic weight removal. The authors validate their approach with a synthetic example and two real‑world applications—voice activity detection and audio anti‑spoofing—providing publicly available code for further research.
By Esteban G\'omez, Tom B\"ackstr\"om
arXiv:2609.10366v1 Announce Type: cross
Abstract: While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true...
By Rishabh Jain, Naomi Harte
arXiv:2606. 19951v1 Announce Type: cross Abstract: Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear.
By Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata