arXiv AI

Assessment in Team Problem-Solving Exercises in Computing Education

arXiv:2607. 19209v1 Announce Type: cross Abstract: This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs).

arXiv Computation and Language
Sep 3

Expos\'ia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.

By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
arXiv Machine Learning
Sep 14

Simulating Disengaged Students to Evaluate LLM-based Tutors

The paper introduces Disengagement-Aware Student Simulators (DAS2), a protocol that models five learner-engagement states—engaged, gaming, wheel-spinning, off-task, and mixed—to evaluate AI tutor performance before deployment. Using annotated tutoring sessions from ASSISTments09, DAS2’s rule-based labels matched human consensus in 81% of cases, and conditioning simulations on intended states narrowed the correctness-rate gap between simulated and authentic sessions for gaming and wheel-spinning behaviors. The study also compares five AI tutors across these states, finding stable relative rankings but state-specific performance differences, and notes that automated evaluation does not fully align with human judgment.

By Xianghui Meng, Jionghao Lin
arXiv AI
Aug 19

WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts

WIP: LLM Odyssey is an open‑source, browser‑based serious gaming platform that offers 13 interactive games to teach Large Language Model engineering concepts such as tokenization, transformer architecture, prompt engineering, RAG, and production deployment. The platform is organized into three learning tiers—Cognitive Core, Systems Forge, and Foundry Arena—aligned with Bloom’s revised taxonomy, and each game employs five pedagogical strategies including immediate feedback, scaffolded hints, progressive difficulty, worked examples, and authentic production scenarios. An initial deployment at a Canadian college in Winter 2026 confirmed functional requirements, highlighted the need for adaptive difficulty, and led to the design of a formal mixed‑methods evaluation protocol for future studies.

By Priyamvada Tripathi