arXiv AI By David Fraile Navarro, Jialei Sheng, Farah Magrabi, Enrico Coiera

Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
2d ago

Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect

The study investigates why large language models (LLMs) show different triage performance when answering clinician‑authored vignettes in multiple‑choice versus free‑text formats. Using sparse‑autoencoder features on Gemma 3 and Qwen3 models, the authors find that medical information is encoded similarly in both formats, but at the decision token the multiple‑choice scaffold dominates, with over 91% of attribution coming from scaffold‑peaking features. The effect varies by model, and shuffling option order eliminates simple positional bias, suggesting the format influence is tied to answer selection rather than earlier case processing.

By David Fraile Navarro, Berardino Como, Jialei Sheng, Soundariya Ananthan, Shlomo Berkovsky
arXiv AI
Aug 3

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.

By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar