arXiv AI

SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Planning

Hugging Face Trending Papers
Jun 19

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.

arXiv Computation and Language
Sep 1

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

The paper introduces Skill-as-Pseudocode (SaP), a method that automatically converts markdown skill libraries for large language model agents into typed pseudocode with deterministic quality control. SaP extracts typed contracts from clusters of procedural passages and verifies them with a four‑check verifier before inlining them into a rewritten skill skeleton that includes both a typed signature and a concrete action template. On the ALFWorld unseen split, SaP outperforms the Graph-of-Skills baseline, achieving 82/402 paired game wins versus 47/402, while reducing input tokens and LLM calls per game.

By Xinze Li, Yuhang Zang, Yixin Cao, Aixin Sun
arXiv AI
Sep 25

Coding Agents for Generalized Task and Motion Planning Problems

The paper investigates whether large language model–based coding agents can automatically synthesize programs that solve generalized task and motion planning (TAMP) problems across diverse instances. Using Claude Code and Codex, the authors evaluate 980 generated programs on 100 held‑out environments from KinDER and PDDLStream, achieving mean success rates between 56 % and 95 %—higher than hand‑engineered planners and other baselines—while requiring an order of magnitude less computation per instance. The study demonstrates that coding agents can calibrate physical models, test edge cases, and refine strategies, suggesting they are a strong baseline for generalized TAMP.

By Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver