Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
arXiv:2606. 18645v1 Announce Type: cross Abstract: Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness.
arXiv:2603. 15988v3 Announce Type: replace-cross Abstract: Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies.
arXiv:2606. 18645v1 Announce Type: cross Abstract: Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness.
The paper evaluates whether audio‑language models can use multimodal clinical context to improve dysarthric speech recognition. Using a benchmark built on the Speech Accessibility Project dataset, the authors test diagnosis labels, clinician ratings, and detailed clinical descriptions as prompts for nine models. They find that these prompts yield negligible or negative effects on word error rate, though fine‑tuning with LoRA and mixed prompt formats reduces WER by 52% and benefits certain subgroups such as Down syndrome and mild‑severity speakers.
arXiv:2602.08696v3 Announce Type: replace-cross Abstract: Dysarthric speech recognition is limited by high speaker variability and scarce labeled data. Existing synthesis methods often couple speaker...
arXiv:2606. 19797v1 Announce Type: cross Abstract: Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria.
The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
The paper evaluates pre‑trained speech embeddings from four state‑of‑the‑art speech foundation models for cross‑lingual Parkinson's disease severity assessment. Experiments span three datasets in zero‑shot and k‑shot settings, showing that these embeddings can transfer meaningfully across languages, though performance varies with dataset characteristics, preprocessing, and adaptation strategy. Misclassifications linked to inter‑speaker variability and atypical speech patterns underscore the need for more robust feature extraction, modeling, and explainability to support reliable clinical insights.
arXiv:2609.21789v1 Announce Type: new Abstract: Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space....
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing ar...
arXiv:2606. 19823v1 Announce Type: cross Abstract: Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability.
BenSparX introduces the first Bengali conversational speech dataset for Parkinson’s disease detection and pairs it with a robust, explainable machine learning framework. The framework uses diverse acoustic features, systematic feature selection, and advanced classifiers, achieving 95.67% accuracy, 95.62% F1, and 0.990 AUC. SHAP analysis is employed to explain feature contributions, and the model outperforms state‑of‑the‑art methods on other language datasets.
The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.