BRACE introduces an anchored Bellman‑residual correction to address stale critic bias in asynchronous reinforcement learning for language models. By limiting the correction horizon to a prefix of policy tokens and adding a constant‑weight Monte‑Carlo tail, it separates policy correction from reward propagation. The method improves mean@1 on BrowseComp‑Plus by 2.4% and runs 2.46× faster per step than synchronous training while staying stable 50 updates off‑policy.
By Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
By Ahad Jawaid
Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.
Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on...
arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.
By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.
By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo