arXiv AI

Structure Enables Effective Self-Localization of Errors in LLMs

arXiv:2602. 02416v2 Announce Type: replace Abstract: Self-correction in language models remains elusive.

arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong
arXiv AI
Sep 3

Stepwise Think-Critique: Interleaved Reasoning and Self-Critique in a Single LLM

The paper introduces Stepwise Think-Critique (STC), an end‑to‑end trainable framework that lets a single large language model (LLM) interleave reasoning with inline, step‑level critique. STC uses reinforcement learning to jointly reward reasoning accuracy and critique consistency, achieving a 7.2% improvement in Pass@1 on five mathematical reasoning benchmarks and a 67.4% step‑level critique F1 score, outperforming several external reward models. This approach aligns LLM behavior more closely with human critical thinking by embedding self‑evaluation directly into the reasoning process.

By Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu