arXiv Computation and Language By Rui Jia, Min Zhang, Fengrui Liu, Bo Jiang, Kun Kuang, Zhongxiang Dai

EduAgentQG: Multi-Agent Personalized Mathematics Question Generation with Explicit Diversity and Objective-Aware Evaluation

Read the original on arXiv Computation and Language →

EduAgentQG is a multi‑agent framework for generating personalized mathematics questions that explicitly controls diversity and aligns with educational objectives. It operates through a closed‑loop cycle of planning, writing, evaluation, refinement, and checking, using fine‑grained evaluation to ensure logical correctness, solvability, and alignment with knowledge concepts, difficulty, grade level, and core competencies. The authors built a benchmark of 10,273 questions across Grades 1‑9 and demonstrated that EduAgentQG outperforms existing methods in diversity, objective consistency, and win rate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 12

Generative AI performance in core undergraduate mathematics: a curriculum-level case study

The study examines how generative AI tools like ChatGPT perform on typical first‑year undergraduate mathematics assessment questions. By generating, transcribing, and blind‑marking AI responses to eight assessments covering the entire curriculum, the authors find that AI attains a first‑class level of performance, with consistency across modules that exceeds that of students in invigilated exams. The results suggest a need to redesign mathematics assessments to address the impact of generative AI.

By Benjamin J. Walker, Nikoleta Kalaydzhieva, Beatriz Navarro Lameda, Ruth A. Reynolds
arXiv AI
Aug 17

TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

arXiv:2608. 13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud