arXiv AI By Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Jul 29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.

By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv Computation and Language
Aug 27

MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

MTDiag is a newly released multi-turn diagnostic dialogue dataset designed to evaluate large language models (LLMs) in clinically meaningful ways. It is built from DDXPlus, MIMIC-IV, and AJCR case reports, covering both common emergency department presentations and rare conditions, and normalizes cases into a canonical schema using UMLS concept identifiers and ICD-10 codes. The dataset includes a UserLM‑8B utterance‑generation pipeline and physician‑validated natural‑language utterances, and introduces clinical knowledge‑grounded metrics that go beyond simple diagnostic accuracy for multi‑turn differential diagnosis tasks.

By Pia Chouayfati, Alexander M. Fichtl, Miriam Ansch\"utz, George Doumat, Georg Groh
arXiv AI
Jul 10

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv:2607. 08257v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters.

By Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, Haoyang Zeng, Jingwang Huang, Kaiwen Wei