arXiv:2605.22286v2 Announce Type: replace-cross
Abstract: Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on...
By Zhaomin Wu, Jiayi Li, Bingsheng He
arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.
By Wenjie Zheng, Qiming Xie, Jianfei Yu, Yang Wang, Lei Cao, Fei Wang, Shijin Wang, Rui Xia, Chengqing Zong
arXiv:2607. 03744v1 Announce Type: new Abstract: Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech.
By Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri, Shrikanth Narayanan
The paper introduces DiaWhisper-DPO, an end‑to‑end model that fine‑tunes Whisper-large-v3 with LoRA and a frame‑level role head to transcribe and attribute utterances in clinical interviews. It further refines the system using failure‑mined preference optimization (DPO) that leverages genuine decoding failures as rejected completions, eliminating the need for human preference data. On the DAIC‑WOZ dataset, DiaWhisper‑DPO attains 0.973 role accuracy and 0.119 DER, outperforming cascaded baselines by 72% and dramatically reducing seed variation, while also improving performance on the cross‑lingual PDCH‑HAMD dataset.
By Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, Jo\~ao Miguel Sanches
arXiv:2608.31007v1 Announce Type: cross
Abstract: Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monito...
By Aowen Shi, Michal Balazia, Danilo Postin, Ren\'e Hurlemann, Jan Alexandersson, Fran\c{c}ois Br\'emond, Philipp M\"uller
arXiv:2606. 10380v1 Announce Type: cross Abstract: Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.
By Grace Byun, Abigail Lott, Rebecca Lipschutz, Sean T. Minton, Elizabeth A. Stinson, Jinho D. Choi
Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re...
arXiv:2609.38480v1 Announce Type: cross
Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients presen...
By Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan
The paper introduces a two‑stage framework for recognizing depression symptoms at the sentence level. First, a contrastively fine‑tuned sentence encoder generates a symptom candidate for each sentence. Then, a fine‑tuned language model verifies the candidate’s presence or absence by comparing the sentence, its context, and a diagnostic definition, ensuring the model’s judgment aligns with that definition before responding.
By Weiming Li, Catarina Barata, Miguel Constante, Joao Sanches
arXiv:2607. 22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability.
By Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
arXiv:2609.38491v1 Announce Type: new
Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...
By Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin
The paper introduces Evidence-Bounded Mental Health Reasoning, addressing the problem that current multimodal mental health screening models treat all clinical speech protocols as equally evidential. It presents the Evidence Package Benchmark, comprising 1,870 annotated packages from six diverse protocols, and proposes EviBound, a protocol-aware framework that limits reasoning to valid evidence using a planner, acoustic consensus, and a boundary critic. EviBound outperforms existing omni-modal baselines, achieving a Depression AUROC of 0.8658 with no claim violations.
By Chengyuan Gao, Jiang Wu, Tao Lu, Jiayan Guo, Mingkun Xu, Tianyi Zang, Shangyang Li