arXiv:2601. 15158v4 Announce Type: replace-cross Abstract: Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought).
By Yuval Ran-Milo, Yotam Alexander, Shahar Mendel, Nadav Cohen
arXiv:2602.13106v2 Announce Type: replace-cross
Abstract: In recent years, there has been growing interest in understanding neural architectures' ability to learn to execute discrete algorithms, a li...
By Solveig Wittig, Antonis Vasileiou, Robert R. Nerem, Timo Stoll, Floris Geerts, Yusu Wang, Christopher Morris
The paper presents an approach to automated theorem proving by framing the construction of clausal connection tableaux as a policy in a transition system. It introduces a graph neural network that scores proof edits based on structure, trained via imitation learning from existing proofs. Experiments on M2k, MPTP2078-bushy, and TPTP v9.2.1 show that the learned policies solve up to 46% more problems than leanCoP and find proofs in an order of magnitude fewer steps.
By Fredrik R{\o}mming, Mantas Bak\v{s}ys, Martin S. Fixman, Sean B. Holden
The paper introduces a novel end‑to‑end, size‑agnostic graph reinforcement learning framework for the one‑dimensional bin packing problem (1D‑BPP). It models packing as a Markov decision process on an item‑compatibility graph, where a graph neural network actor‑critic policy learns to merge compatible partial bins. Empirical results on the BPPLIB benchmark show that the learned policy reduces the mean optimality gap of a constructive heuristic from 2.66 % to 2.31 %, performs competitively against other learned methods, and outperforms a state‑of‑the‑art learned solver on the hardest benchmark family.
By M. Asl{\i} Ayd{\i}n
arXiv:2602. 07832v3 Announce Type: replace-cross Abstract: Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking.
By Xian Wu, Kaijie Zhu, Ying Zhang, Lun Wang, Wenbo Guo
arXiv:2504. 18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}.
By Tianbing Xu