arXiv:2606. 11197v1 Announce Type: cross Abstract: Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings.
By Xuzhi Wang, Xinran Wu, Ziping Zhao, Jianhua Tao, Bj\"orn W. Schuller
arXiv:2607. 21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern.
By Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao
The paper presents a bi‑modal speech‑level transformer that eliminates segment‑level labeling and introduces a hierarchical attention interpretation method. By using gradient‑weighted attention maps from all attention layers, the model provides both speech‑level and sentence‑level explanations of depression detection. Experimental results show the transformer outperforms a segment‑level model (p=0.854 vs. 0.732, r=0.947 vs. 0.808, F1=0.897 vs. 0.768).
By Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia
arXiv:2607. 05901v1 Announce Type: new Abstract: Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries.
By Manning Gao, Tingyi Liu, Leheng Zhang, Haifeng Hu, Yuncheng Jiang, Sijie Mai
MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.
By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.
By Wenjie Zheng, Qiming Xie, Jianfei Yu, Yang Wang, Lei Cao, Fei Wang, Shijin Wang, Rui Xia, Chengqing Zong