Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect
Read the original on arXiv AI →The study investigates why large language models (LLMs) show different triage performance when answering clinician‑authored vignettes in multiple‑choice versus free‑text formats. Using sparse‑autoencoder features on Gemma 3 and Qwen3 models, the authors find that medical information is encoded similarly in both formats, but at the decision token the multiple‑choice scaffold dominates, with over 91% of attribution coming from scaffold‑peaking features. The effect varies by model, and shuffling option order eliminates simple positional bias, suggesting the format influence is tied to answer selection rather than earlier case processing.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.