arXiv AI

Depression Markers in Speech: An Approach based on Tract Variables Dynamics

arXiv:2607. 25888v1 Announce Type: cross Abstract: This study identifies new depression biomarkers based on the dynamical properties of tract variables, which represent geometric features describing the configuration of the speech articulators.

arXiv Computation and Language
Aug 28

Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.

By Jonas L\"anzlinger, Katharina O. E. M\"uller, Burkhard Stiller, Bruno Rodrigues
arXiv AI
Sep 3

Interpretable Symptom Vectors for Depression in a Large Language Model

The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.

By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
arXiv AI
Sep 18

When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews

The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.

By Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galv\'an, Fernando S\'anchez-Vega, A. Pastor L\'opez-Monroy, Petr Motlicek, Esa\'u Villatoro-Tello
arXiv Computation and Language
Sep 21

Hierarchical attention interpretation: an interpretable speech-level transformer for bi-modal depression detection

The paper presents a bi‑modal speech‑level transformer that eliminates segment‑level labeling and introduces a hierarchical attention interpretation method. By using gradient‑weighted attention maps from all attention layers, the model provides both speech‑level and sentence‑level explanations of depression detection. Experimental results show the transformer outperforms a segment‑level model (p=0.854 vs. 0.732, r=0.947 vs. 0.808, F1=0.897 vs. 0.768).

By Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia
arXiv Machine Learning
Aug 31

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

arXiv:2511.07011v2 Announce Type: replace-cross Abstract: Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research t...

By Anastasiia Tokareva, Judith Dineley, Zoe Firth, Pauline Conde, Faith Matcham, Sara Siddi, Femke Lamers, Ewan Carr, Carolin Oetzmann, Daniel Leightley, Yuezhou Zhang, Amos A. Folarin, Josep Maria Haro, Brenda W. J. H. Penninx, Raquel Bailon, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins, The RADAR-CNS Consortium
arXiv Machine Learning
2d ago

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

arXiv:2609.38491v1 Announce Type: new Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...

By Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin
arXiv AI
4d ago

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.

By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv Computation and Language
Sep 14

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.

By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath