arXiv AI

AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks

arXiv AI
3d ago

Consistent Plan-Act for Long-Horizon Agentic Tasks

The paper introduces Consistent Plan-Act (ConPAct), a method that addresses coordination failures between high-level planners and low-level actors in long-horizon agentic tasks. By prompting both agents to produce structured state assertions and programmatically detecting contradictions, the authors identify a systematic planner-actor state mismatch. ConPAct feeds these detected contradictions back to both agents, fine‑tunes them on consistent interactions, and achieves notable performance gains, such as raising MiniGrid success rates from 38.6% to 54.4% with GPT‑5.6‑sol/terra.

By Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye
arXiv AI
Sep 18

Architectural Design, Not Only Model Intelligence, Governs Multi-Agent LLM Performance

The paper argues that the architecture of multi‑agent large language model (LLM) frameworks, rather than just the intelligence of the underlying models, largely determines system performance. It introduces a taxonomy of architectural dimensions—such as orchestration, memory, planning interfaces, specialization, and communication topology—and presents MAFBench, a unified evaluation suite. An empirical study across nine frameworks, keeping the LLM constant, reveals six design principles and shows that choices like orchestration and communication topology can dramatically affect latency, accuracy, and coordination success.

By Abdelghny Orogat, Ana Rostam, Essam Mansour
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Sep 25

Coding Agents for Generalized Task and Motion Planning Problems

The paper investigates whether large language model–based coding agents can automatically synthesize programs that solve generalized task and motion planning (TAMP) problems across diverse instances. Using Claude Code and Codex, the authors evaluate 980 generated programs on 100 held‑out environments from KinDER and PDDLStream, achieving mean success rates between 56 % and 95 %—higher than hand‑engineered planners and other baselines—while requiring an order of magnitude less computation per instance. The study demonstrates that coding agents can calibrate physical models, test edge cases, and refine strategies, suggesting they are a strong baseline for generalized TAMP.

By Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver
Hugging Face Trending Papers
Sep 24

Coding Agents for Generalized Task and Motion Planning Problems

The paper investigates whether coding agents can automate the synthesis of programs that solve generalized Task and Motion Planning (TAMP) problems. By evaluating Claude Code and Codex on 28 simulated environments, the authors find that these agents outperform hand-engineered planners and other baselines, achieving higher success rates and lower computation per instance. The agents also demonstrate adaptive behaviors such as calibrating physical models and refining strategies during interaction.

Hugging Face Trending Papers
Aug 4

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design.

arXiv AI
Aug 20

Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery

Eureka is a task‑conditioned Meta‑Agent architecture that transforms long‑horizon scientific tasks into dynamic obligation graphs with explicit acceptance semantics. During execution it constructs Macro‑Agents equipped with specialized state, memory, operators, tools, verifiers, and local topology, using receding‑horizon planning, architecture promotion, and minimal‑sufficient compilation. The system demonstrates strong empirical performance, completing all 170 recursive tasks, generating 3,948 certificates without false acceptances, and achieving significant reductions in input size, recomputation, and consistent serialization across 16,000 concurrent executions.

By Alizer Wong, Heng Cui, Yi Tan, Xiongchao Zhan, Liang Lin, Yuxiang Guo, Zhaorong Dai, Zixin Zeng, Wenyuan Li