arXiv AI

Performance of a domain-specific large language model in answering patient questions in psychiatry

arXiv AI
Jul 10

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.

By Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, Haoyang Zeng, Jingwang Huang, Kaiwen Wei
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta
arXiv AI
Jul 1

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.

By Mizanur Rahman, Abeer Badawi, Elahe Rahimi, Laleh Seyyed-Kalantari, Frank Rudzicz, Enamul Hoque, Elham Dolatabadi
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker
arXiv AI
Aug 5

Quantifying Hallucinations in Language Language Models on Medical Textbooks

arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.

By Brandon C. Colelough, Davis Bartels, Dina Demner-Fushman