arXiv Machine Learning

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

arXiv:2607. 14171v1 Announce Type: new Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes.

arXiv AI
Sep 18

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG-Tree proposes a compute‑optimal branching strategy for gradient‑efficient reinforcement learning, arguing that branches should be placed where they most reduce policy‑gradient uncertainty per unit of compute. By deriving allocation laws from a law‑of‑total‑variance decomposition, the method introduces an EPIG‑Tree score that guides branch placement using already computed rollouts, estimating occupancy‑ and score‑weighted value uncertainty. Empirical results show EPIG‑Tree reduces gradient MSE in cloned‑state control, improves frozen‑LLM gradient calibration, and outperforms flat GRPO and entropy branching in both single‑turn math and multi‑turn Wordle tasks.

By Nikita Khomich, Leopold Hermansson, Ido Hakimi
arXiv Computation and Language
3d ago

OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

arXiv:2609.40035v1 Announce Type: new Abstract: The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We...

By Junyu Lu, Shichao Weng, Zhiqiang Wang, Haojie Luo, Jingfan Zhang, Yuhua Zhou, Cheng Du, Yuzhuo Zhang, Xi Li, Jinwei Du, Tiancheng Feng, Chuan Xiao, Shuyuan Zheng
arXiv AI
Aug 18

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.

By Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
arXiv AI
Jul 20

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

arXiv:2607. 15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation.

By Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang
arXiv AI
Jun 10

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.

By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
Hugging Face Trending Papers
Sep 10

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by identifying points where a model’s answer belief changes most. Unlike traditional structural or entropy‑based approaches, belief‑shift uses a probe, logit‑lens depth profile, or learned activation direction to locate pivots in the value curve, incurring minimal computational overhead. Experiments across multiple models and benchmarks show that belief‑shift forking consistently outperforms baseline methods, yielding significant gains in mathematics and code tasks.

arXiv AI
Sep 12

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by detecting where a model’s answer belief changes most sharply. Unlike traditional fixed‑length or entropy‑based forking, this approach uses a lightweight probe or learned activation direction to identify pivots in the value curve, reducing unnecessary sampling. Experiments show that belief‑shift forking consistently outperforms baseline methods across multiple models and benchmarks, yielding significant gains in mathematics and code tasks.

By Bin Lei, Yu Li, Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese, Chien-Sheng Wu
arXiv AI
Sep 25

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

The paper introduces GRAFT, a Graph-based Faithful sTep-level credit-assignment framework that constructs a trajectory graph from rollout trajectories, recovers node state-values via Bellman iteration, and assigns step-level advantages based on node value differences. It also proposes Graph GAE to further reduce state-value estimation bias. Experiments on multi-turn agentic benchmarks demonstrate consistent improvements over GRPO and other recent agentic RL algorithms.

By Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang
arXiv AI
Aug 26

Contrastive Branch Policy Optimization

Contrastive Branch Policy Optimization (CBPO) is a reinforcement learning method that separates the allocation of a fixed rollout budget from the translation of branch outcomes into token-level credit. It uses generation entropy to screen branch positions, path- and node-level decay to distribute the budget, and Contrastive Branch Value (CBV) to estimate local decision sensitivity without changing reward signs. CBPO partitions trajectories into non-overlapping credit segments, preventing duplicated gradients and enabling fine-grained credit assignment using only outcome rewards.

By Ying Wang, Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun, Jingli Yang