arXiv AI

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

arXiv:2607. 22996v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes.

arXiv AI
Jul 23

Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

arXiv:2607. 19371v1 Announce Type: new Abstract: Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided inquiry and reveals solutions directly.

By Jing Shao, Qifeng Wu, Hanyu Zhang, Sixia Sun, Jun Zhuang
arXiv AI
Sep 15

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

The paper introduces an exploration-guided prompt scaffolding framework for multimodal large language models, dynamically adjusting the prompt distribution during reinforcement learning post-training. It uses an Exploration Potential Score (EPS) derived from KL-regularized policy improvement to assess prompt utility without extra overhead, and a teacher model rewrites low-utility prompts to preserve intent while improving informativeness. Experiments on Geo3K, MMK12, MathVision, and MMMU-Pro show consistent performance gains, up to 9.7% in-domain and over 11% on out-of-distribution benchmarks.

By Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang
arXiv AI
Aug 26

UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models

The paper introduces UCO, a multi‑turn interactive reinforcement learning method designed to improve adaptive teaching with large language models. UCO employs two reward functions—Progress Reward to gauge genuine cognitive advancement and Scaffold Reward to keep instruction within each student’s Zone of Proximal Development. Experiments on BigMath and MathTutorBench show UCO outperforming 11 baseline models and matching advanced closed‑source systems.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Kun Kuang, Zhongxiang Dai
arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv Machine Learning
Aug 19

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

The paper investigates how different reward specifications affect the reliability of unlearning in large language models using a LoRA-GRPO framework. It compares four reward designs—lexical suppression, anti-refusal shaping, rubric-based broad answering, and explicit refusal contrast—both with and without a supervised fine-tuning warm-up. The results reveal that successful optimization does not guarantee behavioral unlearning, as various evaluation metrics can yield conflicting conclusions due to reward-hacking, policy-support limits, and benchmark probe limitations.

By Rub\'en Balbastre, Juan Manuel Ordu\~na, Mariano P\'erez
arXiv AI
Sep 10

Boosting LLM Reasoning via Human-Inspired Reward Shaping

The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.

By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang