arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.
By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
arXiv:2609.39559v1 Announce Type: new
Abstract: In this work we study the problem of MAPFC, a post-optimization step for Multi-Agent Path Finding (MAPF) plans where we are given a feasible plan produ...
By Oren Salzman
arXiv:2607. 09330v1 Announce Type: new Abstract: Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intelligence such as smart factories, warehouses, and service robotics.
By Nuocheng Yang, Sihua Wang, Zihan Chen, Tony Q. S. Quek, Changchuan Yin
arXiv:2603.08814v2 Announce Type: replace-cross
Abstract: Long-horizon task planning for heterogeneous multi-robot systems is essential for deploying collaborative teams in real-world environments; y...
By Piyush Gupta, Sangjae Bae, Jiachen Li, David Isele
LLM‑BabyBench transforms the BabyAI gridworld into a fully observable, purely textual setting that isolates planning as the sole source of failure. By serialising the entire grid, providing formal instructions, and validating actions deterministically, the benchmark introduces the PPD suite—Predict, Plan, and Decompose tasks—each scored with metrics that separate mission understanding from sequencing. Across a range of large language models, simulation accuracy is high while planning success drops sharply beyond a model‑specific horizon, revealing that plan length—not grid size—drives failure and that models often commit to a single corridor‑shaped route without backtracking.
By Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou
arXiv:2602. 13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models.
By Najmul Hasan, Prashanth BusiReddyGari
GRASP is a multi-stage planning framework that improves the reliability of large language models on complex tasks. It separates planning into three specialized modules—GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation—allowing context isolation and strict macro-regularization. Experiments show GRASP outperforms direct LLM planners by significant margins on datasets such as Natural Plan Calendar Scheduling, ZebraLogic, and SciBench Math, and it mitigates performance collapse in multi-task and dual-task settings.
By Arunabh Srivastava (Amir), Mohammad A. (Amir), Khojastepour, Srimat Chakradhar, Sennur Ulukus
The paper introduces Collective Counterfactual Planning (CCP), a formal model describing how teams coordinate tasks that no single member can handle alone, constrained not by capability but by representational geometry. CCP defines four critical gates—exogenous implementation coalitions, conception, consent, and task-relative verification—that determine whether a team can achieve and legitimately recognize a conjunctive goal. The authors present the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion, and provide a sound and complete four-step solvability scheme under exact representation of relay closure.
By Chainarong Amornbunchornvej
arXiv:2609.22813v1 Announce Type: cross
Abstract: We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. Wh...
By Masafumi Endo, Kohei Honda, Ryo Yonetani
GRASP is a multi-stage planning framework that separates planning into specialized modules: GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation. This strategy-aware approach yields state‑of‑the‑art accuracy on diverse datasets, outperforming direct LLM planners by up to 30.8% on ZebraLogic and reducing multi‑task degradation. GRASP’s context isolation and macro‑regularization also give it a 14.5% edge over frontier reasoning models like GPT‑5‑mini.
The paper introduces Consistent Plan-Act (ConPAct), a method that addresses coordination failures between high-level planners and low-level actors in long-horizon agentic tasks. By prompting both agents to produce structured state assertions and programmatically detecting contradictions, the authors identify a systematic planner-actor state mismatch. ConPAct feeds these detected contradictions back to both agents, fine‑tunes them on consistent interactions, and achieves notable performance gains, such as raising MiniGrid success rates from 38.6% to 54.4% with GPT‑5.6‑sol/terra.
By Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye
arXiv:2607. 17082v2 Announce Type: replace-cross Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results.
By Babak Barazandeh, Subhabrata Majumdar, George Michailidis