arXiv AI

MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultation

arXiv AI
Jul 10

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.

By Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, Haoyang Zeng, Jingwang Huang, Kaiwen Wei
arXiv AI
Jul 16

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.

By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
arXiv Computation and Language
3d ago

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

The paper introduces Evidence-Bounded Mental Health Reasoning, addressing the problem that current multimodal mental health screening models treat all clinical speech protocols as equally evidential. It presents the Evidence Package Benchmark, comprising 1,870 annotated packages from six diverse protocols, and proposes EviBound, a protocol-aware framework that limits reasoning to valid evidence using a planner, acoustic consensus, and a boundary critic. EviBound outperforms existing omni-modal baselines, achieving a Depression AUROC of 0.8658 with no claim violations.

By Chengyuan Gao, Jiang Wu, Tao Lu, Jiayan Guo, Mingkun Xu, Tianyi Zang, Shangyang Li
arXiv Computation and Language
3d ago

An evidence-guided reinforcement learning method to improve psychiatric reasoning in small language models

The paper introduces ClinMPO, an evidence‑guided reinforcement‑learning framework that enhances psychiatric reasoning in small language models (SLMs). ClinMPO leverages a reward model (ClinRM) trained on 18,569 question–answer pairs from 4,474 psychiatry articles and is guided by the psychiatrist‑defined Clinical Psychiatry Thinking Strategy (CPTS). Evaluations on four Qwen3 model sizes show that ClinMPO outperforms baseline, supervised fine‑tuning, and standard policy optimization, with the 8B model surpassing a human baseline of senior pre‑licensure medical students and achieving the highest rank among 31 models. Blinded clinician assessment confirms improved rationale quality across CPTS criteria, demonstrating that clinical evidence and specialist knowledge can be effectively incorporated into medical AI development.

By Xinxin Lin, Guangxin Dai, Yi Zhong, Xiang Li, Xue Xiao, Jian Liu, Yixin Zhang, Lingming Hu, Zhengdong Wu, Yongbo Zheng, Runchuan Zhu, Ming Zhao, Huizi Yu, Yi Zhang, Fangting Lu, Shuo Wu, Jun Zhao, Ping Yin, Joey W. Y. Chan, Ngan Yin Chan, Yumei Wang, Lejin Yang, Yanqiu Xing, Sijing Chen, Yun Kwok Wing, Lin Lu, Xin Ma, Lizhou Fan
arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
arXiv AI
Jul 7

Where do LLMs Fall Short in CBT-Guided Affective Reasoning?

arXiv:2607. 02885v1 Announce Type: cross Abstract: Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction between cognitive and behavioral factors.

By Vaishnavi Sinha, Pooja Guttal, Pranay Deep Reddy Katike, Vishal Sinha, Gerald Ndawula, Lira Yoon, Andrea Kleinsmith, Manas Gaur
arXiv AI
Jul 1

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.

By Mizanur Rahman, Abeer Badawi, Elahe Rahimi, Laleh Seyyed-Kalantari, Frank Rudzicz, Enamul Hoque, Elham Dolatabadi
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam