arXiv AI

Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives

arXiv:2607. 03744v1 Announce Type: new Abstract: Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech.

arXiv AI
Sep 18

When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews

The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.

By Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galv\'an, Fernando S\'anchez-Vega, A. Pastor L\'opez-Monroy, Petr Motlicek, Esa\'u Villatoro-Tello
arXiv AI
4d ago

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.

By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv Computation and Language
Aug 28

Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

The paper proposes a transparent framework that links speech acoustic features—such as pitch variability, pauses, and speech tempo—to DSM‑5 indicators of depression, offering interpretable, indicator‑level outputs instead of opaque black‑box models. It runs locally on commodity hardware to preserve privacy and has been preliminarily evaluated on the DAIC‑WOZ dataset, showing consistent associations between acoustic cues and DSM‑5 indicators of psychomotor change and concentration difficulty. Future work aims to validate the approach on longitudinal data and expand multimodal integration while keeping edge constraints.

By Jonas L\"anzlinger, Katharina O. E. M\"uller, Burkhard Stiller, Bruno Rodrigues
arXiv AI
Jun 30

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

arXiv:2606. 30543v1 Announce Type: cross Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important.

By Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler, Yuzhe Wang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Vel\'azquez
arXiv Computation and Language
Sep 16

DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

The paper introduces DiaWhisper-DPO, an end‑to‑end model that fine‑tunes Whisper-large-v3 with LoRA and a frame‑level role head to transcribe and attribute utterances in clinical interviews. It further refines the system using failure‑mined preference optimization (DPO) that leverages genuine decoding failures as rejected completions, eliminating the need for human preference data. On the DAIC‑WOZ dataset, DiaWhisper‑DPO attains 0.973 role accuracy and 0.119 DER, outperforming cascaded baselines by 72% and dramatically reducing seed variation, while also improving performance on the cross‑lingual PDCH‑HAMD dataset.

By Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, Jo\~ao Miguel Sanches
Hugging Face Trending Papers
Jul 7

Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking

Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries. To address this, we propose a fine-grained multimodal framework featuring a temporal encoder and a mutual transformer to facilitate deep cross-modal fusion.

arXiv Computation and Language
Aug 21

Explainable Multimodal Depression Recognition in Clinical Interviews via PHQ-Aligned Symptom Summarization

arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.

By Wenjie Zheng, Qiming Xie, Jianfei Yu, Yang Wang, Lei Cao, Fei Wang, Shijin Wang, Rui Xia, Chengqing Zong
arXiv Machine Learning
2d ago

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

arXiv:2609.38491v1 Announce Type: new Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...

By Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin