Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks.
arXiv:2607. 18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved.
arXiv:2608.31017v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 1...
The study investigates why large language models (LLMs) show different triage performance when answering clinician‑authored vignettes in multiple‑choice versus free‑text formats. Using sparse‑autoencoder features on Gemma 3 and Qwen3 models, the authors find that medical information is encoded similarly in both formats, but at the decision token the multiple‑choice scaffold dominates, with over 91% of attribution coming from scaffold‑peaking features. The effect varies by model, and shuffling option order eliminates simple positional bias, suggesting the format influence is tied to answer selection rather than earlier case processing.
arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.
arXiv:2608.31016v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...