PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits.
whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."
By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz
arXiv:2608.29397v1 Announce Type: new
Abstract: Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is...
By Zixiang Xu, Jiaan Wang, Fandong Meng
The paper investigates whether large language model–based coding agents can automatically synthesize programs that solve generalized task and motion planning (TAMP) problems across diverse instances. Using Claude Code and Codex, the authors evaluate 980 generated programs on 100 held‑out environments from KinDER and PDDLStream, achieving mean success rates between 56 % and 95 %—higher than hand‑engineered planners and other baselines—while requiring an order of magnitude less computation per instance. The study demonstrates that coding agents can calibrate physical models, test edge cases, and refine strategies, suggesting they are a strong baseline for generalized TAMP.
By Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver
The paper introduces complete cyclic subtask graphs for large language model agents, enabling a workflow controller where all subtasks are fully connected and a unified agent selects transitions based on natural‑language criteria. It evaluates task‑specific and benchmark‑generic cyclic graphs on TextCraft, ALFWorld, and Finance‑Agent, comparing them to ReAct and dependency‑directed workflows, and identifies three distinct workflow signatures that influence the effectiveness of cyclic routing. The study also provides a workflow‑signature matrix, robustness analysis, token‑cost accounting, and failure‑mode structure, concluding that cyclic subtask graphs serve as a diagnostic tool to determine when flexible backtracking is worthwhile versus when simpler controllers suffice.
By Luay Gharzeddine, Samer Saab Jr
The paper introduces InFlowOp, a label‑free optimization framework that assigns costs to each decision in a multi‑agent workflow, balancing agent competence against execution time. It determines task granularity and agent assignment before execution and corrects faults during execution using the same cost metric. The authors also present Braid, a benchmark for multi‑agent coordination, and show that InFlowOp outperforms single‑agent baselines by up to 11.97% across various domains.
By Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng, Qingyun Wang, Haifeng Chen
arXiv:2604. 05150v2 Announce Type: replace-cross Abstract: We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which workflows execute deterministically without further model invocation.
By Geert Trooskens (XY.AI Labs, Palo Alto, CA), Aaron Karlsberg (XY.AI Labs, Palo Alto, CA), Anmol Sharma (XY.AI Labs, Palo Alto, CA), Lamara De Brouwer (XY.AI Labs, Palo Alto, CA), Max Van Puyvelde (Stanford University School of Medicine, Stanford, CA), Matthew Young (XY.AI Labs, Palo Alto, CA), John Thickstun (Cornell University, Ithaca, NY), Gil Alterovitz (Brigham and Women's Hospital / Harvard Medical School, Boston, MA), Walter A. De Brouwer (Stanford University School of Medicine, Stanford, CA)
arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.
By Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung
arXiv:2606. 06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence.
By Ruida Wang, Jerry Huang, Pengcheng Wang, Xuanqing Liu, Luyang Kong, Tong Zhang
The paper introduces Pufibara, an agent harness designed to maintain engineering state and evidence across revisions in Modelica-based physical system modeling. It also presents a 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, evaluated by an external benchmark-owned evaluator. Experiments show Pufibara outperforms Claude Code in task success and resource efficiency across two LLM backends.
By Zizhe Wang
arXiv:2608. 10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures.
By Shuo Hao, You Lu, Bihuan Chen, Xin Peng
The paper investigates whether coding agents can automate the synthesis of programs that solve generalized Task and Motion Planning (TAMP) problems. By evaluating Claude Code and Codex on 28 simulated environments, the authors find that these agents outperform hand-engineered planners and other baselines, achieving higher success rates and lower computation per instance. The agents also demonstrate adaptive behaviors such as calibrating physical models and refining strategies during interaction.
arXiv:2609.38108v1 Announce Type: new
Abstract: Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successfu...
By Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan