StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
Read the original on arXiv AI →StraTA introduces Strategic Trajectory Abstraction, a framework that samples a compact strategy from the initial task state and conditions subsequent actions on that strategy, training strategy generation and action execution jointly with a hierarchical GRPO-style rollout design. The method enhances exploration and credit assignment over long horizons by incorporating diverse strategy rollouts and critical self-judgment. Experiments on ALFWorld, WebShop, and SciWorld demonstrate that StraTA consistently improves sample efficiency and final performance, achieving success rates of 93.1% on ALFWorld, 84.2% on WebShop, and a 63.5% overall score on SciWorld, surpassing frontier closed‑source models.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.