arXiv AI

LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features

arXiv:2606. 28445v1 Announce Type: cross Abstract: Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive screening modality.

arXiv Computation and Language
6d ago

Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring

The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.

By Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li
arXiv AI
Sep 17

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening

MINT (Multimodal Imaging-to-Speech Knowledge Transfer) is a three-stage framework that transfers MRI-derived biomarkers to speech representations for early Alzheimer’s screening. An MRI teacher creates a compact embedding space for CN‑versus‑MCI classification, and a residual projection head aligns speech features to this space using a geometric loss, allowing imaging‑free inference. Experiments on ADNI‑4 show that aligned speech matches speech baselines, while multimodal fusion outperforms MRI alone, and ablations highlight dropout regularization and self‑supervised pretraining as key design choices.

By Vrushank Ahire, Yogesh Kumar, Anouck Girard, M. A. Ganaie
arXiv Machine Learning
Sep 15

CCMAN: Cognitive Instability-Aware Cross-Modal Attention Network for Interpretable Temporal Biomarkers of Verbal Fluency Speech

arXiv:2609.14764v1 Announce Type: cross Abstract: Early detection of cognitive decline from speech offers a scalable and non-invasive alternative to conventional clinical assessment. Verbal fluency t...

By Madhurananda Pahar, Caitlin Illingworth, Dorota Braun, Daniel Blackburn, Heidi Christensen
arXiv Computation and Language
Sep 11

LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

The paper introduces LLM-Anchored Paralinguistic Enrichment (LAPE), a method that enhances language model representations with speech-based paralinguistic cues such as pauses and word elongations. LAPE incorporates three innovations: prosodic event textualization, lexico-prosodic unitization and chunking, and text-anchored paralinguistic fusion using NormGate. Evaluations on the ADReSS and ADReSSo datasets show that LAPE achieves state‑of‑the‑art performance in detecting Alzheimer’s disease from speech.

By Xiao Wei, Yuqin Lin, Yaru Cao, Jinyu Li, Bin Wen, Kai Li, Yueying Chen, Longbiao Wang, Jianwu Dang
arXiv AI
Sep 4

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.

By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso