LeAct: Learning to Reason from Expert Actions
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
arXiv:2608. 14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice.
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
arXiv:2609.01591v1 Announce Type: new Abstract: AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for...
arXiv:2608. 12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment.
arXiv:2607. 16097v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it.
arXiv:2608.27757v1 Announce Type: new Abstract: Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc...
The paper presents a method for generating creative chess puzzles using masked diffusion models that can be conditioned on tactical themes and partial board positions. It introduces an auxiliary best‑move prediction task that boosts solution uniqueness by 11.6% and theme‑conditioning accuracy by 2.5%. A reinforcement learning framework further increases the yield of unique, theme‑matching puzzles by 89.1%, and the authors release open‑weights models for the community.
The paper presents a tutoring platform that combines a generative AI chatbot with a reinforcement learning algorithm to adaptively sequence practice problems for students learning Python. In a five‑month field study across ten high schools, the adaptive sequencing improved unassisted final exam performance by 0.15 standard deviations, with mediation analysis indicating that higher engagement drove the gains. The study demonstrates that signals from student‑chatbot interactions can be leveraged to personalize and optimize learning at scale.
arXiv:2511.05933v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is often credited with improving reasoning at the expense of factual knowledge. We instead find that reasoning mo...
arXiv:1808. 07645v5 Announce Type: replace-cross Abstract: The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity.
Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.
arXiv:2601. 18778v3 Announce Type: replace Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal.
arXiv:2607. 28638v1 Announce Type: cross Abstract: As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights.