arXiv AI

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.

arXiv AI
Jul 1

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.

By Mizanur Rahman, Abeer Badawi, Elahe Rahimi, Laleh Seyyed-Kalantari, Frank Rudzicz, Enamul Hoque, Elham Dolatabadi
arXiv AI
Sep 21

Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

The paper introduces a clinician‑grounded evaluation platform called InterviewPlayground, which uses a memory‑augmented patient simulator to assess AI‑assisted psychiatric intake systems. It supports comparison across different interviewing styles, reduces clinician workload, and measures clinically relevant performance. In a pilot study, a GPT‑based intake interviewer captured more relevant items but made more unfounded inferences and missed safety concerns compared to clinicians.

By King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob Taylor, Ananya Joshi
arXiv AI
Jun 17

AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows

arXiv:2606. 17474v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based, limiting their ability to reflect the sequential, uncertain, and interactive nature of real-world care.

By Jiahui Niu, Huizi Yu, Wenkong Wang, Guangxin Dai, Jingxian He, Xiang Li, Zhiying Liang, Xinxin Lin, Kent CY So, Bryan YP Yan, Yun Kwok Wing, Yanqiu Xing, Xin Ma, Lizhou Fan
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv Computation and Language
Aug 27

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

HealthBench-Psych is a mental‑health subset of the OpenAI HealthBench benchmark, created by filtering 5,000 physician‑rubric conversations for mental‑health relevance using an LLM‑applied rubric and validating the selection through two rounds of blinded clinician review. The resulting 610 conversations (12.2 % of the corpus) are released along with a pipeline, model responses, grades, and analysis code. Evaluation of 20 frontier and open models by a cross‑vendor panel of three LLM judges shows a statistically tied frontier cluster, measurable refusal behavior in two models, and near‑identical rankings across judges (τ ≥ 0.92).

By Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous
arXiv AI
Sep 15

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

arXiv:2609.15855v1 Announce Type: cross Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...

By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
arXiv AI
Jul 16

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.

By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler