arXiv:2607. 13389v1 Announce Type: new Abstract: Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget.
By Patrick Wilhelm, Odej Kao
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.
By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv:2610.09679v1 Announce Type: new
Abstract: Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to...
By Yiming Zong, Yige Wang, Xing Hu, Jiashuo Jiang, Zuo-Jun Max Shen
The paper proposes a method for selectively querying language‑model advice in reinforcement learning by predicting the value of potential responses and only querying when the expected benefit outweighs the cost. It introduces a certified, response‑contingent metareasoning framework that guarantees near‑optimal advice usage under certain assumptions, and demonstrates that a calibrated controller with Qwen2.5 advisors can improve task performance while drastically reducing the number of advice calls on the BabyAI benchmark.
By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan
The paper introduces LP‑BTS, a learning‑guided planning framework for mobile charging in large, dynamic action spaces. It uses a graph proposal policy to narrow candidate stops, a value critic to evaluate leaf nodes, and edge‑budgeted PUCT to compare short simulated futures before action selection. Experiments on a 30‑scenario battery‑life benchmark show LP‑BTS achieving the highest survival and alive‑AUC, outperforming domain‑engineered baselines and heuristic policies.
By Liang-Ching Tao, Pi-Chung Wang
arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.
By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv:2606. 19236v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training.
By Haipeng Luo, Qingfeng Sun, Songli Wu, Can Xu, Wenfeng Deng, Han Hu, Yansong Tang
arXiv:2606. 05606v1 Announce Type: new Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt, despite large differences in the training signal different prompts provide.
By Yiming Zong, Yige Wang, Jiashuo Jiang
arXiv:2608.21501v1 Announce Type: new
Abstract: Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: evidence about success, a transport oper...
By Qifan Shi, Zhaolu Kang, Chenghua Zhu
arXiv:2606. 08779v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses.
By Jiashun Liu, Runze Liu, Xu Wan, Jing Liang, Hongyao Tang, Ling Pan
arXiv:2608. 01130v1 Announce Type: new Abstract: A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward.
By Yuyang Shen