arXiv Machine Learning
Aug 27

Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

The paper introduces Reflection Steering, a training‑free framework that separates reflection-related activations from general reasoning in large language models. By contrasting reflective and non‑reflective hidden states, denoising with PCA, and orthogonalizing against reasoning directions, the method selectively removes reflection computation while preserving accuracy. Experiments on two benchmarks and three open‑weight LLMs show an average 16.9% reduction in reasoning tokens, and a tunable parameter α allows deployment‑time trade‑offs between token savings, accuracy, and stability.

By Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang
arXiv AI
Aug 5

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

arXiv:2608. 03972v1 Announce Type: new Abstract: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models.

By Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Hugging Face Trending Papers
Aug 4

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples.

arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong