GRASP is a multi-stage planning framework that separates planning into specialized modules: GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation. This strategy-aware approach yields state‑of‑the‑art accuracy on diverse datasets, outperforming direct LLM planners by up to 30.8% on ZebraLogic and reducing multi‑task degradation. GRASP’s context isolation and macro‑regularization also give it a 14.5% edge over frontier reasoning models like GPT‑5‑mini.
arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.
By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv:2607. 08894v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational costs and stochastic behavior.
By Maureese Williams, Dymitr Nowicki
arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.
By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.
arXiv:2606. 29932v1 Announce Type: new Abstract: Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward.
By Tianyu Jin, Shuo Chen, Yida Wang, Liuyu Xiang, Yingzhuo Liu, Zhiyao Jiang, Yexin Li, Zhaofeng He
arXiv:2606. 27806v1 Announce Type: new Abstract: World models for language agents come in two useful forms.
By Xinyuan Song, Zekun Cai
arXiv:2603.08814v2 Announce Type: replace-cross
Abstract: Long-horizon task planning for heterogeneous multi-robot systems is essential for deploying collaborative teams in real-world environments; y...
By Piyush Gupta, Sangjae Bae, Jiachen Li, David Isele
The paper introduces CLUE, a framework that lets robots actively resolve contextual uncertainty for underspecified natural language tasks. CLUE employs an LLM-derived policy to generate task-relevant hypotheses and plans, then uses an online language-embedded map to ground these into actions, refining its plan through closed-loop interaction. Experiments on a Boston Dynamics Spot across diverse indoor and outdoor settings show CLUE achieving near-oracle performance and outperforming LLM planners without closed-loop feedback by a significant margin.
By Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar
The paper investigates whether large language model–based coding agents can automatically synthesize programs that solve generalized task and motion planning (TAMP) problems across diverse instances. Using Claude Code and Codex, the authors evaluate 980 generated programs on 100 held‑out environments from KinDER and PDDLStream, achieving mean success rates between 56 % and 95 %—higher than hand‑engineered planners and other baselines—while requiring an order of magnitude less computation per instance. The study demonstrates that coding agents can calibrate physical models, test edge cases, and refine strategies, suggesting they are a strong baseline for generalized TAMP.
By Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver
PlanPO introduces a group planning-aware policy optimization method for multi-turn agentic large language models, addressing the issue of advantage collapse caused by treating all successful trajectories equally. By incorporating coarse-to-fine advantage signals that reflect differences in trajectory and turn lengths, PlanPO encourages agents to learn generalizable planning and generation behaviors. Experiments show a 27.2% average improvement over GRPO on benchmarks such as ALFWorld, WebShop, and SciWorld, with minimal extra training cost.
By Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu