MisEdu‑RAG is a dual‑hypergraph retrieval‑augmented generation framework designed to help novice math teachers diagnose and remediate student misconceptions. It structures pedagogical knowledge as a concept hypergraph and real student mistake cases as an instance hypergraph, performing two‑stage retrieval to ground responses in both layers. On the MisstepMath dataset, MisEdu‑RAG outperforms baseline models, improving token‑F1 by 10.95% and achieving up to 15.3% higher quality across five dimensions, especially in diversity and empowerment.
By Zhihan Guo, Yuting Lu, Jionghao Lin
arXiv:2508. 17092v2 Announce Type: replace-cross Abstract: Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content.
By Yahya Badran, Christine Preisach
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
The paper investigates how large language models (LLMs) generate distractor answers for multiple‑choice questions (MCQs) by modeling student misconceptions. It introduces a learning‑science‑based taxonomy of reasoning strategies and applies it to LLM‑generated reasoning traces in math and science MCQs. The study finds that in math, LLMs often follow a misconception‑based process that can be diagnostically useful, whereas in science they rely more on semantic similarity, with frequent failures when the model cannot produce a correct solution or discards plausible distractors. Providing the correct solution in the prompt improves alignment with human distractors by 6.4%.
"whyItMatters":"The findings show that anchoring distractor generation to the correct solution enhances LLM alignment with human‑authored distractors, underscoring the importance of correct‑answer cues in educational AI."
By Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan
IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.
By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
arXiv:2605.30051v2 Announce Type: replace
Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as stude...
By Zhangqi Duan, Shuyan Huang, Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew Lan