arXiv:2607. 19371v1 Announce Type: new Abstract: Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided inquiry and reveals solutions directly.
By Jing Shao, Qifeng Wu, Hanyu Zhang, Sixia Sun, Jun Zhuang
arXiv:2606. 09887v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness.
By Zirui Liu, Jie Ouyang, Qi Liu, Xianquan Wang, Jiayu Liu, Tingyue Pan, Qingchuan Li, Jing Sha, Zhenya Huang, Shijin Wang, Enhong Chen
arXiv:2606. 11744v1 Announce Type: cross Abstract: Large language models are now widely used for everyday learning, but the underlying interactions are typically unstructured chats rather than following a curriculum.
By Sidney Tio, Arunesh Sinha, Pradeep Varakantham
The paper introduces an exploration-guided prompt scaffolding framework for multimodal large language models, dynamically adjusting the prompt distribution during reinforcement learning post-training. It uses an Exploration Potential Score (EPS) derived from KL-regularized policy improvement to assess prompt utility without extra overhead, and a teacher model rewrites low-utility prompts to preserve intent while improving informativeness. Experiments on Geo3K, MMK12, MathVision, and MMMU-Pro show consistent performance gains, up to 9.7% in-domain and over 11% on out-of-distribution benchmarks.
By Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang
The paper introduces UCO, a multi‑turn interactive reinforcement learning method designed to improve adaptive teaching with large language models. UCO employs two reward functions—Progress Reward to gauge genuine cognitive advancement and Scaffold Reward to keep instruction within each student’s Zone of Proximal Development. Experiments on BigMath and MathTutorBench show UCO outperforming 11 baseline models and matching advanced closed‑source systems.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Kun Kuang, Zhongxiang Dai
arXiv:2606. 18257v1 Announce Type: cross Abstract: While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied.
By Xiaolong Wang, Zhe Zhao, Song Lai, Chaoli Zhang, Zijie Geng, Yu Tong, Ye Wei, Qingsong Wen
EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.
By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv:2609.24290v1 Announce Type: new
Abstract: Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confident...
By Yunxiang Li, Xixin Wu, Helen Meng
The paper investigates how different reward specifications affect the reliability of unlearning in large language models using a LoRA-GRPO framework. It compares four reward designs—lexical suppression, anti-refusal shaping, rubric-based broad answering, and explicit refusal contrast—both with and without a supervised fine-tuning warm-up. The results reveal that successful optimization does not guarantee behavioral unlearning, as various evaluation metrics can yield conflicting conclusions due to reward-hacking, policy-support limits, and benchmark probe limitations.
By Rub\'en Balbastre, Juan Manuel Ordu\~na, Mariano P\'erez
The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.
By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
arXiv:2606. 09052v1 Announce Type: cross Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision.
By Siyu Chen, Miao Lu, Beining Wu, Heejune Sheen, Fengzhuo Zhang, Shuangning Li, Zhiyuan Li, Jose Blanchet, Tianhao Wang, Zhuoran Yang
arXiv:2608. 03952v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners.
By Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen