arXiv AI By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov

MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution

Read the original on arXiv AI →

arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
arXiv Computation and Language
Sep 21

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.

By Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
arXiv AI
6d ago

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

StraTA introduces Strategic Trajectory Abstraction, a framework that samples a compact strategy from the initial task state and conditions subsequent actions on that strategy, training strategy generation and action execution jointly with a hierarchical GRPO-style rollout design. The method enhances exploration and credit assignment over long horizons by incorporating diverse strategy rollouts and critical self-judgment. Experiments on ALFWorld, WebShop, and SciWorld demonstrate that StraTA consistently improves sample efficiency and final performance, achieving success rates of 93.1% on ALFWorld, 84.2% on WebShop, and a 63.5% overall score on SciWorld, surpassing frontier closed‑source models.

By Xiangyuan Xue, Yifan Zhou, Zidong Wang, Shengji Tang, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin