arXiv Machine Learning By Joshua Mitton, Prarthana Bhattacharyya, Digory Smith, Thomas Christie, Ralph Abboud, Simon Woodhead

Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank

Read the original on arXiv Machine Learning →

arXiv:2602. 02414v2 Announce Type: replace-cross Abstract: Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 11

MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math Teachers

MisEdu‑RAG is a dual‑hypergraph retrieval‑augmented generation framework designed to help novice math teachers diagnose and remediate student misconceptions. It structures pedagogical knowledge as a concept hypergraph and real student mistake cases as an instance hypergraph, performing two‑stage retrieval to ground responses in both layers. On the MisstepMath dataset, MisEdu‑RAG outperforms baseline models, improving token‑F1 by 10.95% and achieving up to 15.3% higher quality across five dimensions, especially in diversity and empowerment.

By Zhihan Guo, Yuting Lu, Jionghao Lin
arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv AI
Sep 16

Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation

The paper investigates how large language models (LLMs) generate distractor answers for multiple‑choice questions (MCQs) by modeling student misconceptions. It introduces a learning‑science‑based taxonomy of reasoning strategies and applies it to LLM‑generated reasoning traces in math and science MCQs. The study finds that in math, LLMs often follow a misconception‑based process that can be diagnostically useful, whereas in science they rely more on semantic similarity, with frequent failures when the model cannot produce a correct solution or discards plausible distractors. Providing the correct solution in the prompt improves alignment with human distractors by 6.4%. "whyItMatters":"The findings show that anchoring distractor generation to the correct solution enhances LLM alignment with human‑authored distractors, underscoring the importance of correct‑answer cues in educational AI."

By Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan
arXiv Computation and Language
Aug 27

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.

By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky