HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims to provide a shared standard for model performance and safety in health.
arXiv:2511.11689v4 Announce Type: replace-cross
Abstract: Generative AI chatbots built for mental health could extend access to care, but evidence from real-world use is limited. We report a single-a...
By Thomas D. Hull, Lizhe Zhang, Caitlin A. Stamatis, Patricia A. Arean, Matteo Malgaroli
arXiv:2603. 25821v2 Announce Type: replace-cross Abstract: We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physician-patient interactions.
By Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk
The paper introduces CounselReflect, a tool that converts counseling quality metrics into a framework for users to reflect on their mental‑health AI conversations. Through interviews with 21 users, the study finds that while most participants rarely reflect on their interactions, they identify specific questions they would like such a tool to address. The findings also reveal that users tend to confirm existing beliefs and focus on familiar dimensions, highlighting the need for reflection tools to expose blind spots and encourage a more comprehensive examination of AI interactions, especially when revisiting emotionally charged exchanges.
By Yahan Li, Chaohao Du, Christopher Chun Kuizon, Zeyang Li, Nimra Ishfaq, Shupeng Cheng, Angelica Yinling Sun, Adam C. Frank, Angel Hsing-Chi Hwang, Ruishan Liu
arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.
By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.
By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv:2609.15855v1 Announce Type: cross
Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...
By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
arXiv:2507. 02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering.
By Mohammad Anas Azeez, Rafiq Ali, Ebad Shabbir, Zohaib Hasan Siddiqui, Gautam Siddharth Kashyap, Jiechao Gao, Usman Naseem
The paper introduces a clinician‑grounded evaluation platform called InterviewPlayground, which uses a memory‑augmented patient simulator to assess AI‑assisted psychiatric intake systems. It supports comparison across different interviewing styles, reduces clinician workload, and measures clinically relevant performance. In a pilot study, a GPT‑based intake interviewer captured more relevant items but made more unfounded inferences and missed safety concerns compared to clinicians.
By King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob Taylor, Ananya Joshi
arXiv:2603. 25821v3 Announce Type: replace-cross Abstract: We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions.
By Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk, Andy Nkansah
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
By Anabela C. Areias, Catarina Botelho, Ant\'onio Farinhas, Areti Vassilopoulos, Dora Janela, Xin Tong, Nuno M. Guerreiro, Maya D'Eon, Fab\'iola Costa, Ricardo Rei
The paper surveys 61 studies on mental‑health AI and identifies a misalignment in how trust is evaluated across disciplines. It proposes a three‑layer framework—human‑oriented, interaction‑oriented, and AI‑oriented trust—and maps stakeholder perspectives onto these layers. The authors argue that future research should focus on calibrating human trust to actual interaction and AI trustworthiness rather than merely maximizing perceived trust.
By Xin Sun, Yue Su, Yifan Mo, Qingyu Meng, Yuxuan Li, Min Chen, Mengyuan Zhang, Saku Sugawara, Charlotte Gerritsen, Sander L. Koole, Koen Hindriks, Jiahuan Pei