The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.
By Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing ar...
The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.
By Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li
The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and time-consuming nature of language-specific model training. We propose a novel solution using cross-language training to detect AD in languages beyond those used for model training.
The paper investigates the interpretability of Domain-Adapted Prompt-based Fine-tuning (DAPF) models for dementia detection from spoken language. Using probing and analysis techniques, the authors find that DAPF achieves strong overall performance (accuracy = 0.83, macro‑F1 = 0.83) and that the diagnosis is most recoverable from its [MASK] representation. However, token‑level explanations derived from DAPF are largely driven by language‑task vocabulary, discourse markers, and transcription artifacts, and perturbation tests reveal weak or negative effects, indicating that the masked‑token interface captures diagnosis information without providing faithful token‑level explanations.
By Pardis Ranjbar-Noiey, Natalie Parde
arXiv:2603. 13673v2 Announce Type: replace Abstract: Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for early-stage detection and disease staging.
By Mingchen Shao, Yuzhang Xie, Carl Yang, Jiaying Lu
arXiv:2607. 10168v1 Announce Type: cross Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available.
By Rashin Gholijani Farahani, Azam Bastanfard
The paper introduces XTF, an explainable token‑level noise filtering framework for fine‑tuning large language models. XTF breaks down token contributions into reasoning importance, knowledge novelty, and task relevance, scores them, and masks gradients of noisy tokens to improve fine‑tuning. Experiments on math, code, and medicine tasks across seven LLMs show up to a 13.7% performance boost over standard fine‑tuning.
By Yuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu, Hongbin Zhou, Lan Tao, Yiming Li, Zhan Qin, Kui Ren
The paper introduces LLM-Anchored Paralinguistic Enrichment (LAPE), a method that enhances language model representations with speech-based paralinguistic cues such as pauses and word elongations. LAPE incorporates three innovations: prosodic event textualization, lexico-prosodic unitization and chunking, and text-anchored paralinguistic fusion using NormGate. Evaluations on the ADReSS and ADReSSo datasets show that LAPE achieves state‑of‑the‑art performance in detecting Alzheimer’s disease from speech.
By Xiao Wei, Yuqin Lin, Yaru Cao, Jinyu Li, Bin Wen, Kai Li, Yueying Chen, Longbiao Wang, Jianwu Dang
The paper introduces a unified pre‑training framework for medical representations that incorporates hierarchical sub‑token aggregation, partial masking, and cross‑reference mechanisms to better capture the structure of medical codes. The resulting model outperforms existing BERT‑based approaches on pre‑training tasks and downstream clinical predictions, such as dementia onset and hospitalization. An in‑silico drug repositioning study for Alzheimer’s disease demonstrates the framework’s ability to rediscover known drugs and prioritize new hypotheses without external literature, establishing a workflow for hypothesis generation and prioritization based on observational data.
By Yuhei Fujioka, Daitaro Misawa, Shingo Fukuma
This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.
By Rakib Abdullah, Md. Maruful Islam Maruf
arXiv:2606. 30675v1 Announce Type: cross Abstract: Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging.
By Olivier Jiyoun Jung, Jonghyeon Park, Myungwoo Oh