arXiv Machine Learning

CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models

CanvasAnneal is a curriculum‑guided reinforcement learning framework designed to improve Diffusion Language Models (DLMs) on complex reasoning and tool‑use tasks. It starts training by injecting reasoning traces from a stronger teacher model into the diffusion canvas, then gradually reduces this guidance so the model learns to generate reasoning independently. Experiments on mathematical reasoning and tool‑use benchmarks show that CanvasAnneal outperforms standard diffusion RL methods such as diffu‑GRPO on tasks like MATH500, Countdown, and Tau2, and accelerates reward improvement, though the gains vary by task.

arXiv AI
Sep 10

Boosting LLM Reasoning via Human-Inspired Reward Shaping

The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.

By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
Hugging Face Trending Papers
Aug 3

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation

Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration. New methods are required that leverage the broad knowledge and flexibility of pre-trained LLMs to deliberately generate diverse experience at training time.

arXiv Machine Learning
Aug 4

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation

arXiv:2608. 02087v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration.

By Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
arXiv AI
Aug 3

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.

By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
arXiv Computation and Language
Aug 28

SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning

SPEAR (Symbolic Process Evaluation and Alignment Reward) is a training‑free, plug‑and‑play reward method for on‑policy distillation in reinforcement learning. It converts natural‑language reasoning traces into domain‑adaptive symbolic milestones and uses the longest common subsequence to align student exploration with teacher milestones, producing a dense, order‑aware reward that enforces logical consistency without an external neural verifier. Experiments on math, science, and commonsense reasoning tasks show that SPEAR effectively bridges the reasoning gap between student and teacher models through sequence‑level distillation with efficient dense process rewards.

By Zhuochun Li, Yuelyu Ji, Yiming Zeng, Daqing He
arXiv AI
2d ago

SPIRAL: Learning to Search and Aggregate

SPIRAL is a reinforcement‑learning framework that trains language models to employ three inference primitives—sequential reasoning within a trace, parallel sampling of independent traces, and aggregation of those traces—within a single compute pipeline. The model first generates multiple independent chain‑of‑thought traces in parallel, then produces a final aggregation trace conditioned on them, with all components optimized end‑to‑end for the reward of the aggregated response. Experiments on reasoning tasks demonstrate that SPIRAL scales efficiently with inference compute, achieving up to 11× better scaling efficiency and 15% higher performance compared to the GRPO baseline when all three primitives are scaled.

By Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman
Hugging Face Trending Papers
Jun 22

SPIRAL: Learning to Search and Aggregate

Language model reasoning can be substantially improved at test time via scaffolds that scale inference compute across different primitives -- sequential reasoning within a trace, independently sampled parallel traces, and aggregation of multiple reasoning traces into a final response. During post-training, however, language models are optimized only for sequential reasoning within a single trace.

arXiv Machine Learning
Sep 23

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.

By Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang
arXiv AI
Aug 6

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

arXiv:2608. 02087v2 Announce Type: replace Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration.

By Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
arXiv AI
Jul 28

Offline-Online Curriculum RL for Multimodal Reasoning

arXiv:2607. 23700v1 Announce Type: new Abstract: Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers.

By Wendi Deng, Hang Du, Guoshun Nan, Haokun Tian, Jiaqi Yu, Xinlei Cao, Jaile Li, Jingfeng Chen, Ling Deng, Ting Li, Hao Yang, Jun Liu, Xudong Jiang, Sicong Leng