arXiv:2607. 25679v1 Announce Type: cross Abstract: Multimodal behavioral analysis offers a scalable approach to assessing depression, anxiety, and stress, yet generic fusion models often ignore the psychometric structure of questionnaire labels.
By Shiyu Teng, Haichen Yu, Jiaqing Liu, Hao Sun, Yu Song, Shurong Chai, Ruibo Hou, Lanfen Lin, Yen-Wei Chen
arXiv:2609.39049v1 Announce Type: cross
Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
By Xinkai Chen
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easie...
arXiv:2606. 03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized.
By Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
arXiv:2608.21868v1 Announce Type: new
Abstract: Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile....
By Ao Chen, Xiaojiang Peng
BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
LongCounsel-8 is a new benchmark suite comprising three datasets with 7,749 five‑session counseling dialogues, each grounded in real client profiles, depression trajectories, symptom compositions, and counseling patterns. The benchmark addresses challenges of longitudinal consistency, empirical grounding of symptom progression, and natural expression of controlled depression states without exposing labels. Experiments show that lower single‑session error does not ensure accurate trend detection, methods perform worse on worsening trajectories, and adding more session history can reduce trend prediction accuracy.
By Jiayi Li, Zhaomin Wu, Bingsheng He
The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.
By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
The study introduces a clinician‑in‑the‑loop benchmark to assess whether large language models can generate evidence‑grounded Brief Hierarchical Taxonomy of Psychopathology (B‑HiTOP) item profiles from multimodal data, including passive sensing, ecological momentary assessment, and questionnaires. Using the GLOBEM dataset, the authors create 14,592 participant‑day instances aligned to 29 B‑HiTOP items across five spectra, and evaluate evidence compatibility rather than diagnostic accuracy. Two‑stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, yielding more conservative score distributions across models, spectra, and evidence settings.
By Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
arXiv:2605.22286v2 Announce Type: replace-cross
Abstract: Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on...
By Zhaomin Wu, Jiayi Li, Bingsheng He
The paper introduces the first benchmark for evaluating confidence estimation in large language models during multi‑turn medical consultations, combining three types of medical data and an information sufficiency gradient to capture how confidence and correctness evolve as evidence accumulates. Experiments with 27 methods reveal that token‑level and consistency‑level confidence approaches are limited by medical data, and that medical reasoning must be judged on both diagnostic accuracy and information completeness. Building on these findings, the authors propose MedConf, a retrieval‑augmented, linguistically grounded self‑assessment framework that aligns patient information with supporting, missing, and contradictory relations, producing interpretable confidence estimates that outperform existing methods across multiple datasets and LLMs.
By Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao
The paper introduces Expected‑Severity‑Risk (ESR), a new objective for selecting questions in proactive medical dialogue that prioritizes reducing the expected severity of diagnostic errors rather than merely uncertainty. ESR uses population statistics to marginalize over possible answers and distills its rankings into a prefix‑only language policy, enabling deployment without teacher‑side risk computation. Experiments on DDxPlus show ESR cuts high‑severity diagnostic misses by 29.5% and boosts accuracy while adding only 0.14 extra questions per dialogue.
By Chenxuan Li, Xinrong Chen, Luyan Zhang, Peidong Jia, Runfan Zheng, Zhongyu Zhao, Xuecheng Shang, Peixing Wan