arXiv:2607. 22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability.
By Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.
By Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galv\'an, Fernando S\'anchez-Vega, A. Pastor L\'opez-Monroy, Petr Motlicek, Esa\'u Villatoro-Tello
MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.
By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.
By Jonas L\"anzlinger, Katharina O. E. M\"uller, Burkhard Stiller, Bruno Rodrigues
arXiv:2606. 30543v1 Announce Type: cross Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important.
By Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler, Yuzhe Wang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Vel\'azquez
arXiv:2606. 11197v1 Announce Type: cross Abstract: Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings.
By Xuzhi Wang, Xinran Wu, Ziping Zhao, Jianhua Tao, Bj\"orn W. Schuller
The paper introduces DiaWhisper-DPO, an end‑to‑end model that fine‑tunes Whisper-large-v3 with LoRA and a frame‑level role head to transcribe and attribute utterances in clinical interviews. It further refines the system using failure‑mined preference optimization (DPO) that leverages genuine decoding failures as rejected completions, eliminating the need for human preference data. On the DAIC‑WOZ dataset, DiaWhisper‑DPO attains 0.973 role accuracy and 0.119 DER, outperforming cascaded baselines by 72% and dramatically reducing seed variation, while also improving performance on the cross‑lingual PDCH‑HAMD dataset.
By Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, Jo\~ao Miguel Sanches
Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries. To address this, we propose a fine-grained multimodal framework featuring a temporal encoder and a mutual transformer to facilitate deep cross-modal fusion.
arXiv:2607. 05901v1 Announce Type: new Abstract: Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries.
By Manning Gao, Tingyi Liu, Leheng Zhang, Haifeng Hu, Yuncheng Jiang, Sijie Mai
arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.
By Wenjie Zheng, Qiming Xie, Jianfei Yu, Yang Wang, Lei Cao, Fei Wang, Shijin Wang, Rui Xia, Chengqing Zong
arXiv:2605.22286v2 Announce Type: replace-cross
Abstract: Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on...
By Zhaomin Wu, Jiayi Li, Bingsheng He
arXiv:2609.38491v1 Announce Type: new
Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...
By Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin