arXiv AI

Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering

arXiv:2606. 16890v1 Announce Type: cross Abstract: Aggregate accuracy benchmarks conceal a systematic structure in how large language models fail at electronic health record (EHR) question answering: questions requiring more inferential steps produce disproportionately more errors.

arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Aug 24

Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing

The paper identifies a new problem in clinical natural language processing called the clinical lost‑in‑the‑middle (CLitM) effect, where large language models perform poorly on information located near the center of long electronic health record (EHR) documents. Using the MedAlign dataset, the authors quantify a 21.9‑percentage‑point accuracy gap across 2,196 instruction‑response pairs and six models, showing that most critical facts lie in the CLitM trough. They propose Query‑Conditioned Clinical Suppression (QCCS), a lightweight context‑selection gate that outperforms traditional retrieval methods (BM25, dense retrieval, cross‑encoder reranking) on a held‑out set of 83 instructions, achieving up to 25.3% accuracy for middle‑position queries. whyItMatters":"The study demonstrates that standard retrieval strategies fail to reliably surface central clinical information, and that a query‑aligned selection mechanism can substantially improve model performance on critical EHR data."

By Sanjay Basu
arXiv AI
Aug 13

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.

By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji
arXiv AI
2d ago

Scaling Clinical Judgment to Evaluate Medical AI

arXiv:2609.12822v2 Announce Type: replace Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....

By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker
arXiv AI
2d ago

Jev in Medicine: A Benchmark Evaluation

The study evaluates Jev 1.13, a non‑generative model that selects from predefined answer options, on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena‑MCQ, and the NEJM Case Challenges. Jev’s top‑1 accuracy matches GPT‑6 Sol with medium reasoning on PubMedQA but falls behind on MetaMedQA, DiagnosisArena‑MCQ, and NEJM cases. While Jev shows strong calibration on MetaMedQA and is fast and inexpensive, its performance on examination and complex diagnostic tasks is substantially lower, indicating the need for task‑specific validation before clinical deployment.

By Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho