arXiv Machine Learning By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

Read the original on arXiv Machine Learning →

arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv AI
Aug 5

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).

By Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon