arXiv AI

Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect

The study investigates why large language models (LLMs) show different triage performance when answering clinician‑authored vignettes in multiple‑choice versus free‑text formats. Using sparse‑autoencoder features on Gemma 3 and Qwen3 models, the authors find that medical information is encoded similarly in both formats, but at the decision token the multiple‑choice scaffold dominates, with over 91% of attribution coming from scaffold‑peaking features. The effect varies by model, and shuffling option order eliminates simple positional bias, suggesting the format influence is tied to answer selection rather than earlier case processing.

arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv Computation and Language
Sep 25

Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark

The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.

By Alexander Apartsin, Yehudit Aperstein
arXiv Machine Learning
2d ago

Clinical Concept Centers in LLMs

The paper investigates whether clinical concepts are represented as distinct, causally influential centers within the latent space of large language models (LLMs). By evaluating eleven open-weight LLMs, the authors discover that each model contains dedicated clinical concept centers that are interpretable, activate only on relevant clinical narratives, and drive model behavior in both constrained and open-ended contexts. These centers can be leveraged for evaluation and performance improvement, as steering models along them enhances downstream clinical outcomes and aligns with clinician preferences.

By Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker
arXiv AI
5d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv AI
Jun 30

Primary ICD Category Prediction using LLM-based Probing

arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.

By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen