arXiv AI By Jaesung Bae, Xiuwen Zheng, Minje Kim, Chang D. Yoo, Mark Hasegawa-Johnson

Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech

Read the original on arXiv AI →

arXiv:2603. 15988v3 Announce Type: replace-cross Abstract: Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

The paper evaluates whether audio‑language models can use multimodal clinical context to improve dysarthric speech recognition. Using a benchmark built on the Speech Accessibility Project dataset, the authors test diagnosis labels, clinician ratings, and detailed clinical descriptions as prompts for nine models. They find that these prompts yield negligible or negative effects on word error rate, though fine‑tuning with LoRA and mixed prompt formats reduces WER by 52% and benefits certain subgroups such as Down syndrome and mild‑severity speakers.

By Pehu\'en Moure, Niclas Pokel, Bilal Bounajma, Yingqiang Gao, Roman Boehringer, Longbiao Cheng, Shih-Chii Liu
arXiv Computation and Language
Sep 15

Controllable Dysarthric Speech Synthesis with Patient-Specific Conditioning for Speaker-Diverse ASR Augmentation

arXiv:2602.08696v3 Announce Type: replace-cross Abstract: Dysarthric speech recognition is limited by high speaker variability and scarce labeled data. Existing synthesis methods often couple speaker...

By Haoshen Wang, Xueli Zhong, Bingbing Lin, Jia Huang, Xingduo Pan, Shengxiang Liang, Nizhuan Wang, Wai Ting Siok
arXiv AI
Sep 2

Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.

By Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman