The paper investigates how large language models (LLMs) generate distractor answers for multiple‑choice questions (MCQs) by modeling student misconceptions. It introduces a learning‑science‑based taxonomy of reasoning strategies and applies it to LLM‑generated reasoning traces in math and science MCQs. The study finds that in math, LLMs often follow a misconception‑based process that can be diagnostically useful, whereas in science they rely more on semantic similarity, with frequent failures when the model cannot produce a correct solution or discards plausible distractors. Providing the correct solution in the prompt improves alignment with human distractors by 6.4%.
"whyItMatters":"The findings show that anchoring distractor generation to the correct solution enhances LLM alignment with human‑authored distractors, underscoring the importance of correct‑answer cues in educational AI."
By Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan
The paper proposes a response‑free method for estimating difficulty of reading‑comprehension multiple‑choice items by fine‑tuning a transformer on item wording. It introduces two extensions to a baseline joint‑encoding model: a component‑wise variant that encodes passage, question, and options separately, and a multi‑task variant that adds a question‑answering auxiliary task. Experiments on a corpus of nearly 30,000 items show that both extensions outperform the baseline, especially the multi‑task variant across all metrics and the component‑wise variant in rank ordering, even with limited training data.
By Jan Net\'ik, Patr\'icia Martinkov\'a
arXiv:2606. 28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction.
By Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
MUDDLE is a benchmark designed to disentangle the effects of document length and topical distractors on document question‑answering systems. It contains 270 human‑annotated questions, each tested in five conditions: the source alone, the source with two or four hard negatives (topically similar), and the source with two or four random distractors matched in length and provenance. Experiments with GPT‑5‑mini show that hard negatives reduce accuracy more than length‑matched random distractors, indicating that topical similarity is a more significant source of error than length alone.
By Jason Luo, Saibilila Abudukelimu, Judy Song, Andrew Feng, Shivank Garg, Vasu Sharma, Kevin Zhu
arXiv:2501. 06286v2 Announce Type: replace-cross Abstract: Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models remain vulnerable to irrelevant context.
By Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
arXiv:2511. 21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation.
By Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach