arXiv Machine Learning

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.

arXiv AI
2d ago

Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

The paper introduces Independent–Communicate–Revise (ICR), a framework that isolates communication effects in large language model multi‑agent systems by fixing initial reasoning and measuring how messages influence answer revision. ICR evaluates correction, preservation, and selectivity across four reasoning benchmarks, revealing that similar overall accuracy can mask divergent revision behaviors. The study shows that richer messages can both improve and harm outcomes, and that receiver policies can shift preservation and correction dynamics differently across tasks.

By Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan
Hugging Face Trending Papers
5d ago

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

The paper explores Retrospection-Only Fine-Tuning (ROFT), a method where a language-model agent improves its behavior by generating and training on explanations of its own experiences, without external teachers or reward signals. In software‑engineering tasks with Qwen3.5‑4B, ROFT achieves comparable or better solve rates than GRPO while requiring fewer updates and training time, and can learn from failures alone. Behavioral analysis shows ROFT indirectly assigns credit to actions and can produce shorter, more direct solutions when prompted to focus on direct solutions.

arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong
arXiv AI
4d ago

Principled Thoughts for Latent Recursive LLM Systems

The paper introduces REST (REpresentation‑Supervised Thoughts), a new training objective for latent recursive language‑model systems that supplements cross‑entropy loss with differentiable penalties enforcing causality, minimality, separability, and stability of internal thought representations. By applying REST to both single‑agent and multi‑agent setups without changing architectures or adding inference parameters, the authors achieve up to 7.5 percentage‑point gains in accuracy across seven diverse benchmarks and a 30 % improvement in convergence to the final answer. The method also yields more informative latent thoughts, improving interpretability of agent communication.

By Fahd Seddik, Fatemeh Fard
arXiv AI
Jul 15

Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents

arXiv:2607. 12397v1 Announce Type: new Abstract: LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed.

By Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin