The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.
By Jonas L\"anzlinger, Katharina O. E. M\"uller, Burkhard Stiller, Bruno Rodrigues
arXiv:2606. 05561v1 Announce Type: cross Abstract: Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns about demographic information exposure.
By Xueyang Wu, Siyuan Liu, Kezhuo Yang, Guang Ling
The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.
By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
arXiv:2607. 03744v1 Announce Type: new Abstract: Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech.
By Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri, Shrikanth Narayanan
arXiv:2607. 00986v1 Announce Type: new Abstract: Automatically detecting stress in speech provides an unobtrusive way to gain insights relevant to behavioral research or clinical assessment.
By Hanna Drimalla, Wieland R. Cremer, Christine Kraus, Oliver T. Wolf
arXiv:2607. 22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability.
By Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.
By Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galv\'an, Fernando S\'anchez-Vega, A. Pastor L\'opez-Monroy, Petr Motlicek, Esa\'u Villatoro-Tello
The paper presents a bi‑modal speech‑level transformer that eliminates segment‑level labeling and introduces a hierarchical attention interpretation method. By using gradient‑weighted attention maps from all attention layers, the model provides both speech‑level and sentence‑level explanations of depression detection. Experimental results show the transformer outperforms a segment‑level model (p=0.854 vs. 0.732, r=0.947 vs. 0.808, F1=0.897 vs. 0.768).
By Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia
arXiv:2511.07011v2 Announce Type: replace-cross
Abstract: Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research t...
By Anastasiia Tokareva, Judith Dineley, Zoe Firth, Pauline Conde, Faith Matcham, Sara Siddi, Femke Lamers, Ewan Carr, Carolin Oetzmann, Daniel Leightley, Yuezhou Zhang, Amos A. Folarin, Josep Maria Haro, Brenda W. J. H. Penninx, Raquel Bailon, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins, The RADAR-CNS Consortium
arXiv:2609.38491v1 Announce Type: new
Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...
By Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin
MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.
By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.
By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath