arXiv:2608. 07378v1 Announce Type: cross Abstract: Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes.
By Xin Wang, Yingchao Huang, Yuhan Su, Shanshan Yao, Wei Peng
The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.
By Yunchong Xiao, Yuxiang Zhao, Ziyang Ma, Shuai Wang, Kai Yu, Jiachun Liao, Xie Chen
The paper investigates on-device language models (ODLMs) for predicting stress in a mobile health context, focusing on privacy-preserving, cloud-independent inference. Using zero‑shot prompting, the authors evaluate ODLMs across multimodal data—objective sensor features and subjective self‑reports—measuring predictive accuracy, latency, and throughput. Results indicate that sensor features slightly outperform self‑reports, and that lightweight sub‑2B models deliver low latency with predictable resource usage, underscoring both the potential and practical limits of ODLMs for mobile mental health.
By Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker
The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.
By Jonas L\"anzlinger, Katharina O. E. M\"uller, Burkhard Stiller, Bruno Rodrigues
arXiv:2607. 21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern.
By Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao
arXiv:2607. 22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability.
By Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
arXiv:2607. 25888v1 Announce Type: cross Abstract: This study identifies new depression biomarkers based on the dynamical properties of tract variables, which represent geometric features describing the configuration of the speech articulators.
By Sahar Altalhi, Tanaya Guha, Alessandro Vinciarelli
arXiv:2606. 18571v1 Announce Type: new Abstract: Mild Cognitive Impairment (MCI) is a medical condition characterized by a noticeable decline in memory, language, or thinking abilities.
By William Nguyen, Jiali Cheng, Hadi Amiri
The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.
By Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galv\'an, Fernando S\'anchez-Vega, A. Pastor L\'opez-Monroy, Petr Motlicek, Esa\'u Villatoro-Tello
arXiv:2607. 03744v1 Announce Type: new Abstract: Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech.
By Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri, Shrikanth Narayanan
The paper examines why speech‑based screening for Alzheimer’s disease fails to generalize across different languages, tasks, and recording protocols. Using a leave‑one‑corpus‑out evaluation on four datasets, it finds that 59 of 70 interpretable speech features show conflicting patterns between healthy controls and cognitive risk groups, with pause, silence, and speech rate being highly protocol‑sensitive. The authors propose a fusion method that combines XLM‑R text baseline scores with evidence anchors, improving mean speaker AUC to 0.785 and worst‑case AUC to 0.615, and emphasize the importance of auditing feature transferability and reporting worst‑case domain robustness.
By Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li
The paper introduces VoxPrivacy, a benchmark for assessing interactional privacy in Speech Language Models (SLMs). It evaluates models on a 32‑hour bilingual dataset across three difficulty tiers, revealing that most open‑source SLMs perform near random on conditional privacy decisions and even strong closed‑source systems struggle with proactive privacy inference. The authors also validate these findings on a real‑speech subset and show that fine‑tuning on a 4,000‑hour training set can improve privacy‑preserving capabilities while maintaining robustness.
By Yuxiang Wang, Hongyu Liu, Dekun Chen, Xueyao Zhang, Zhizheng Wu