arXiv Machine Learning

Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

The study investigates whether acoustic factors encoded in pretrained self‑supervised learning (SSL) models can systematically influence predictions in speech‑based Alzheimer's disease (AD) assessment. Using the ADReSSo dataset and three large SSL backbones, the authors applied controlled noise and reverberation interventions to various audio segments and combined layer‑wise decoding, input‑ and representation‑space interventions, and geometric alignment analysis. The results demonstrate that such acoustic interventions alter AD predictions across all models, with noise producing the strongest effect, and that these effects are structured relative to the classifier’s decision direction and reproducible on a held‑out test set.

arXiv AI
Sep 2

Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.

By Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman
arXiv Computation and Language
6d ago

Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring

The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.

By Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li
arXiv Machine Learning
Jul 28

Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender

arXiv:2607. 23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human listeners is underexplored across languages and genders, where pathological markers and perceptual strategies differ.

By Liu He, Yuanchao Li, Yin-Long Liu, Rui Feng, Yiming Wang, Jiaxin Chen, Yizhe Wang, Jiahong Yuan
arXiv AI
3d ago

Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling

The study introduces an automated system for estimating the Verbal Fluency Index (VFI) in individuals with Motor Neuron Disease (MND) by combining ASR (WhisperX) and VAD (Silero) with precise timestamping. Using a unique MND dataset, the approach outperformed traditional acoustic and self‑supervised embedding methods, achieving high predictive accuracy (R² up to 0.9 for P‑words and 0.8 for S‑words). Clinically inspired features were consistently superior, demonstrating the feasibility of automated VFI estimation for monitoring cognitive impairment in MND.

By Bahman Mirheidari, Leslie Ing, Daniel Blackburn, Sharon Abrahams, Christopher McDermott, Heidi Christensen
arXiv Machine Learning
Aug 18

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

arXiv:2608. 15037v1 Announce Type: cross Abstract: Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference.

By Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
arXiv AI
Jun 30

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.

By Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke