arXiv Computation and Language

Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.

arXiv AI
4d ago

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.

By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv AI
Sep 16

EviDep: Uncertainty-Aware Multimodal Depression Estimation via Disentangled Evidential Learning

EviDep is a multimodal evidential regression framework for estimating depression severity from audio–visual recordings, incorporating multi‑scale temporal modeling and shared–private representation learning. It uses frequency‑aware feature extraction to decompose behavioral sequences into multiple frequency bands, refined by scale‑specific experts, and applies disentangled evidential learning to separate cross‑modal shared and modality‑specific information. The model outputs Normal‑Inverse‑Gamma distributions via multi‑branch evidential regression, enabling estimation of depression severity along with aleatoric and epistemic uncertainty, and demonstrates competitive accuracy on several benchmark datasets.

By Fangyuan Liu, Sirui Zhao, Yangsong Zhang, Jinyang Huang, Feng-Qi Cui, Bin Luo, Tong Xu, Enhong Chen
arXiv Computation and Language
Aug 21

Explainable Multimodal Depression Recognition in Clinical Interviews via PHQ-Aligned Symptom Summarization

arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.

By Wenjie Zheng, Qiming Xie, Jianfei Yu, Yang Wang, Lei Cao, Fei Wang, Shijin Wang, Rui Xia, Chengqing Zong
arXiv AI
Sep 3

Interpretable Symptom Vectors for Depression in a Large Language Model

The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.

By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
Hugging Face Trending Papers
Jul 7

Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking

Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries. To address this, we propose a fine-grained multimodal framework featuring a temporal encoder and a mutual transformer to facilitate deep cross-modal fusion.

arXiv Computation and Language
Sep 21

Hierarchical attention interpretation: an interpretable speech-level transformer for bi-modal depression detection

The paper presents a bi‑modal speech‑level transformer that eliminates segment‑level labeling and introduces a hierarchical attention interpretation method. By using gradient‑weighted attention maps from all attention layers, the model provides both speech‑level and sentence‑level explanations of depression detection. Experimental results show the transformer outperforms a segment‑level model (p=0.854 vs. 0.732, r=0.947 vs. 0.808, F1=0.897 vs. 0.768).

By Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia