arXiv AI

Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies

The paper evaluates strategies for increasing plan‑execution flexibility by converting sequential plans into partial‑order plans through deordering and reordering. It compares block deordering methods, which restructure causal dependencies, with MaxSAT‑based approaches that optimize within existing causal structures. The study finds that block deordering consistently outperforms MaxSAT in both effectiveness and efficiency, offering anytime solutions and higher flexibility gains per computation time.

arXiv AI
Sep 4

Lose the Order, Keep the Hierarchy: Deordering HTN Plans

The paper "Lose the Order, Keep the Hierarchy: Deordering HTN Plans" adapts two classical plan deordering techniques to the Hierarchical Task Network (HTN) planning framework, extending them to respect hierarchical decomposition constraints. The authors evaluate their methods on the IPC 2023 Partial-Order HTN benchmarks and compare them with Optiplan, an HTN planner that generates partially ordered plans directly. Results show a substantial reduction in ordering constraints, with a smaller but noticeable decrease in critical path length.

By Takudzwa Togarepi, Gaspard Quenard, Damien Pellier, Humbert Fiorino
arXiv AI
Sep 7

Substrate-Aware AI Agents: Execution Context as a First-Class Input

The paper introduces the concept of substrate blindness, where AI agents lack execution context in their planning. By providing a 128 MB RAM and 10 s wall‑time contract to large language models, the authors show that agents generate code that uses less memory, runs faster, and incorporates structural changes such as bounded blocking and in‑place buffers. Across three leading models, contract disclosure improved resource usage and correctness, demonstrating that minimal execution contracts can guide agents to produce more efficient programs.

By Manu Agrawal
arXiv AI
Aug 26

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.

By Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Jul 21

CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents

arXiv:2607. 16632v1 Announce Type: cross Abstract: Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuous, tool feedback is delayed and heterogeneous, and a backend failure may require revising RTL rather than tuning another physical-design parameter.

By Peilong Zhou, Zhirong Chen, Cangyuan Li, Haoyu Gao, Kaiyan Chang, Ziming Qu, Ying Wang