AHEAD is a step‑aware framework that augments reinforcement learning for multi‑turn LLM agents by matching different supervision sources to different step types. The teacher receives environment feedback on all steps and LLM‑generated corrective hints only on error steps, providing finer‑grained guidance than uniform trajectory‑level rewards. Across ALFWorld, WebShop, and Search‑based QA, AHEAD improves task success by 13.3 points on ALFWorld and 11.0 on WebShop at 7B, reaches target success rates faster, and solves tasks within tighter interaction budgets compared to outcome‑only RL and prior self‑distillation baselines.
By Xiaolong Jin, Dingmin Wang, Vijay Lingam, Varun Kumar
HINT-SD introduces a targeted self‑distillation framework for long‑horizon language‑model agents that uses full‑trajectory hindsight to identify failure‑relevant actions and applies feedback‑conditioned distillation only to those action spans. This selective approach reduces the need for per‑turn feedback, improving training efficiency and effectiveness. Experiments on BFCL v3 and AppWorld demonstrate that HINT‑SD outperforms dense per‑turn feedback baselines by up to 13.60 percentage points on average while cutting training time per step by 2.26×.
By Woongyeong Yeo, Yumin Choi, Taekyung Ki, Sung Ju Hwang
arXiv:2608. 01597v1 Announce Type: new Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed.
By Haowei Liu, Jiamian Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang
arXiv:2609.36864v1 Announce Type: new
Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling com...
By Fanchao Chen, Hengyu Fu, Shivaram Venkataraman, Jiantao Jiao
arXiv:2608. 04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions.
By Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu
arXiv:2606. 11559v1 Announce Type: new Abstract: Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns.
By Haoran Liu, Yuwei Zhang, Xiyao Li, Bohan Lyu, Jingbo Shang