arXiv AI

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

arXiv:2608. 11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs).

arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong
arXiv AI
Jun 2

CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards

arXiv:2606. 00020v1 Announce Type: cross Abstract: Large Language Model (LLM) based Chinese Grammatical Error Correction (CGEC) systems face two critical challenges: general-purpose models lack specialized linguistic priors for subtle grammatical distinctions, and Supervised Fine-Tuning (SFT) with Maximum Likelihood Estimation fails to optimize for precision-focused metrics, leading to systematic over-correction.

By Wei Tian, Yuhao Zhou, Man Lan
arXiv AI
Jul 7

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

arXiv:2607. 05199v1 Announce Type: new Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows.

By Raj Jaiswal, Dhruv Jain, Rishabh Dhawan, Sree Krishna Uppalapati, Shin'ichi Satoh, Tanuja Ganu, Rajiv Ratn Shah
arXiv Computation and Language
Sep 2

From Rollouts to Recipes: Self-Contained Post-Training for LLMs

The paper introduces Self‑Routing, a post‑training framework that tailors optimization for each sample based on its rollout correctness and confidence. Instead of applying a single recipe to all data, samples are routed to different strategies—GRPO, on‑policy self‑distillation, regularization, or skipped—allowing training to adapt without external teachers or extra annotations. Experiments on Qwen3 and Qwen3.5 show consistent improvements over uniform methods and reveal that the routing distribution evolves during training, reducing unnecessary updates on low‑signal or already stable samples.

By Yifei Li, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Jun Liu