arXiv AI By Yuval Ran-Milo, Yotam Alexander, Shahar Mendel, Nadav Cohen

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data

Read the original on arXiv AI →

arXiv:2601. 15158v4 Announce Type: replace-cross Abstract: Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 31

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

arXiv:2607. 26119v1 Announce Type: new Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear.

By Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani
Hugging Face Trending Papers
Jun 22

Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reasoning performance at inference time. In this work, we theoretically analyze why reinforcement fine-tuning induces better reasoning ability than purely supervised fine-tuning (SFT) methods.

arXiv AI
Sep 18

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

The paper introduces a dependency‑graph framework to formalize compositional reasoning in language models, defining three increasing levels of compositionality. Using data‑structure tasks with deterministic rewards, the authors observe a consistent asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks transfers more readily to decomposed ones. They provide a theoretical explanation for this asymmetry and evaluate its effects under length extrapolation, structural distribution shift, and transfer to unseen skills, concluding with a pilot study on real‑world tool‑calling benchmarks that suggests the phenomenon extends to practical settings.

By Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik