Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
arXiv:2608. 07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards.
arXiv:2608. 01597v1 Announce Type: new Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed.
arXiv:2608. 07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards.
AHEAD is a step‑aware framework that augments reinforcement learning for multi‑turn LLM agents by matching different supervision sources to different step types. The teacher receives environment feedback on all steps and LLM‑generated corrective hints only on error steps, providing finer‑grained guidance than uniform trajectory‑level rewards. Across ALFWorld, WebShop, and Search‑based QA, AHEAD improves task success by 13.3 points on ALFWorld and 11.0 on WebShop at 7B, reaches target success rates faster, and solves tasks within tighter interaction budgets compared to outcome‑only RL and prior self‑distillation baselines.
arXiv:2608. 12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment.
HINT-SD introduces a targeted self‑distillation framework for long‑horizon language‑model agents that uses full‑trajectory hindsight to identify failure‑relevant actions and applies feedback‑conditioned distillation only to those action spans. This selective approach reduces the need for per‑turn feedback, improving training efficiency and effectiveness. Experiments on BFCL v3 and AppWorld demonstrate that HINT‑SD outperforms dense per‑turn feedback baselines by up to 13.60 percentage points on average while cutting training time per step by 2.26×.
arXiv:2607. 23955v1 Announce Type: new Abstract: Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior.
arXiv:2608. 05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer.
arXiv:2610.00838v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error...
arXiv:2607. 11172v1 Announce Type: new Abstract: Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage.
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approache...
arXiv:2607. 23263v1 Announce Type: new Abstract: Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning.
The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.
arXiv:2609.36864v1 Announce Type: new Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling com...