Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
arXiv:2607. 21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern.
The study introduces an automated system for estimating the Verbal Fluency Index (VFI) in individuals with Motor Neuron Disease (MND) by combining ASR (WhisperX) and VAD (Silero) with precise timestamping. Using a unique MND dataset, the approach outperformed traditional acoustic and self‑supervised embedding methods, achieving high predictive accuracy (R² up to 0.9 for P‑words and 0.8 for S‑words). Clinically inspired features were consistently superior, demonstrating the feasibility of automated VFI estimation for monitoring cognitive impairment in MND.
arXiv:2607. 21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern.
arXiv:2606. 17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems.
arXiv:2607. 10168v1 Announce Type: cross Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available.
The study investigates whether acoustic factors encoded in pretrained self‑supervised learning (SSL) models can systematically influence predictions in speech‑based Alzheimer's disease (AD) assessment. Using the ADReSSo dataset and three large SSL backbones, the authors applied controlled noise and reverberation interventions to various audio segments and combined layer‑wise decoding, input‑ and representation‑space interventions, and geometric alignment analysis. The results demonstrate that such acoustic interventions alter AD predictions across all models, with noise producing the strongest effect, and that these effects are structured relative to the classifier’s decision direction and reproducible on a held‑out test set.
arXiv:2606. 30675v1 Announce Type: cross Abstract: Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging.
arXiv:2505. 23378v3 Announce Type: replace Abstract: Speaker-dependent modelling can substantially improve performance in speech-based health monitoring applications.
The paper evaluates pre‑trained speech embeddings from four state‑of‑the‑art speech foundation models for cross‑lingual Parkinson's disease severity assessment. Experiments span three datasets in zero‑shot and k‑shot settings, showing that these embeddings can transfer meaningfully across languages, though performance varies with dataset characteristics, preprocessing, and adaptation strategy. Misclassifications linked to inter‑speaker variability and atypical speech patterns underscore the need for more robust feature extraction, modeling, and explainability to support reliable clinical insights.
arXiv:2502. 03484v3 Announce Type: replace-cross Abstract: Dementia encompasses a group of syndromes that impair cognitive functions such as memory, reasoning, and the ability to perform daily activities.
arXiv:2606. 14788v1 Announce Type: cross Abstract: Voice-based screening offers a scalable and non-invasive way to assess neurodegenerative diseases such as Alzheimer's disease (AD) and Parkinson's disease (PD), but their staging remains challenging due to the difficulty of integrating heterogeneous data.
arXiv:2606. 16505v1 Announce Type: cross Abstract: Understanding speaker confidence is crucial in educational settings, as it can enhance personalised feedback and improve learning outcomes.
BenSparX introduces the first Bengali conversational speech dataset for Parkinson’s disease detection and pairs it with a robust, explainable machine learning framework. The framework uses diverse acoustic features, systematic feature selection, and advanced classifiers, achieving 95.67% accuracy, 95.62% F1, and 0.990 AUC. SHAP analysis is employed to explain feature contributions, and the model outperforms state‑of‑the‑art methods on other language datasets.
The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.