arXiv AI By Holger Maus, Fabian Kieser, Stefan Petersen, Peter Wulff, Paul Tschisgale

Developing an LLM-Based Feedback System Grounded in Evidence-Centered Design to Support Physics Problem Solving

Read the original on arXiv AI →

arXiv:2512. 10785v3 Announce Type: replace-cross Abstract: Generative AI offers new opportunities for individualized and adaptive learning, e.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Feedback Without the Wait: Piloting a Generative AI Practice Platform in a Large Maths Class

The paper describes a pilot study of a generative AI practice platform designed to give immediate, scaffolded feedback to students in a large mathematics class. By moving human oversight to the verification of solutions rather than real‑time grading, the platform aims to reduce delays while maintaining trust and accountability. The study investigates student engagement, perceived value, and reliability of the AI feedback, and explores how these findings might apply to other engineering subjects.

By Lili Chen, Gavin Buskes, Yuxin Ren, Chin Tong Leong
arXiv AI
Sep 7

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

The paper introduces a verifier‑guided explainable reasoning framework for educational question answering that integrates gold‑anchored QLoRA, a task‑aware symbolic router, and group‑relative RLVR. It adapts Qwen2.5‑3B‑Instruct with field‑weighted QLoRA supervision, routes logic problems to a FOL/Z3 verifier and physics problems to a symbolic solver, and uses verifier feedback for candidate evaluation, self‑revision, and reward construction. Experiments on 438 held‑out examples show that RLVR boosts reasoning depth (P3) from 50.68 % to 72.20 %, while symbolic verification improves answer reliability at the system level.

By Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran, Duy Phuong Tran
arXiv AI
Sep 24

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).

By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang