arXiv Computation and Language
The study evaluates seven parameter‑efficient fine‑tuning (PEFT) methods—LoRA, QLoRA, AdaLoRA, DoRA, LoHA, VeRA, and VB‑LoRA—on two ASR back‑ends (Whisper‑large‑v3 and Qwen3‑ASR‑1.7B) for a single post‑stroke Hungarian male speaker with severe dysarthria. Attention‑projection adapters consistently lower character error rates (CER) on both models, with LoRA emerging as the simplest and most effective choice; QLoRA performs worse and offers no memory advantage at this scale. Full fine‑tuning yields the lowest CER, but a 115 MB LoRA that also adapts feed‑forward blocks achieves comparable accuracy with only 3.7 % of the per‑patient storage, and a 5‑minute enrollment audio captures nearly half of the zero‑shot‑to‑30‑minute CER improvement.
whyItMatters":"The paper demonstrates that lightweight PEFT adapters can substantially improve dysarthric ASR performance while keeping storage and computational costs low, offering a practical path for personalized speech recognition in clinical settings."
arXiv:2606. 19797v1 Announce Type: cross Abstract: Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria.
By Paban Sapkota, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shrikanth Narayanan
The paper evaluates whether audio‑language models can use multimodal clinical context to improve dysarthric speech recognition. Using a benchmark built on the Speech Accessibility Project dataset, the authors test diagnosis labels, clinician ratings, and detailed clinical descriptions as prompts for nine models. They find that these prompts yield negligible or negative effects on word error rate, though fine‑tuning with LoRA and mixed prompt formats reduces WER by 52% and benefits certain subgroups such as Down syndrome and mild‑severity speakers.
By Pehu\'en Moure, Niclas Pokel, Bilal Bounajma, Yingqiang Gao, Roman Boehringer, Longbiao Cheng, Shih-Chii Liu
arXiv:2607. 08256v1 Announce Type: cross Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier.
By Taehyung Yu, Seongjae Kang
arXiv:2607. 07985v1 Announce Type: cross Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.
By A. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan
The paper presents an interpretable, fair, and accurately benchmarked automated system for assessing second‑language English speaking. Using a hybrid of feature‑based speech‑timing metrics and a large language model (LLM) fluency judgment, the system achieves a Spearman correlation of 0.818 with the ICNALE Global Rating Archive, outperforming 81 % of trained human raters. A controlled study shows that encoding pauses into the LLM prompt does not meaningfully affect fluency scores, indicating that the system’s fluency signal derives from measurable speech‑timing features.
By Eichi Uehara
arXiv:2607. 26410v1 Announce Type: cross Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.
By Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg
arXiv:2605. 00865v2 Announce Type: replace-cross Abstract: We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark.
By Xiaoyang Li, Zeyan Tao
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory. md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best.
SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.
By Biraj Subedi
arXiv:2606. 07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standard German subtitles as weak supervision.
By Felix Akeret
arXiv:2607. 26472v1 Announce Type: cross Abstract: Audio deepfake detectors often degrade when generators, corpora, or recording conditions change.
By Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei, Xinhai Chen, Xiangyang Wang, Yigui Feng, Shuai Li, Gencheng Liu, Keqi Yang, Qinglin Wang
arXiv:2603. 04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription.
By Akif Islam, Raufun Nahar, Md. Ekramul Hamid