Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
arXiv:2608. 07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards.
arXiv:2608. 07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards.
The paper introduces ActObs, a supervised fine‑tuning method that, unlike standard approaches, also predicts environment observations in agent trajectories. While both ActObs and action‑only training perform similarly after initial fine‑tuning, ActObs diverges during subsequent reinforcement learning, yielding higher pass@k scores on several benchmarks and better cross‑domain task performance. The authors attribute this advantage to ActObs’s joint supervision, which preserves observation gradients and prevents the policy from over‑specializing on actions alone.
arXiv:2607. 14171v1 Announce Type: new Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes.
arXiv:2607. 09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern.
arXiv:2607. 15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation.
arXiv:2608. 01418v1 Announce Type: cross Abstract: Autoregressive rollout generation is a major computational cost in reinforcement learning for large language models.
Contrastive Branch Policy Optimization (CBPO) is a reinforcement learning method that separates the allocation of a fixed rollout budget from the translation of branch outcomes into token-level credit. It uses generation entropy to screen branch positions, path- and node-level decay to distribute the budget, and Contrastive Branch Value (CBV) to estimate local decision sensitivity without changing reward signs. CBPO partitions trajectories into non-overlapping credit segments, preventing duplicated gradients and enabling fine-grained credit assignment using only outcome rewards.
arXiv:2605.28295v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollout...
arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...
arXiv:2608. 19842v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models.
arXiv:2607. 08837v1 Announce Type: cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers.