arXiv AI

LLM-Only PDDL Domain Repair with Open-Weight Models

The paper evaluates how well open-weight large language models can repair Planning Domain Definition Language (PDDL) models using only LLMs. Experiments show that while the best LLM achieves an F1 score of 0.87—an improvement of 0.38 over a symbolic baseline—it still fails to reliably satisfy test constraints, with a mean test pass rate of only 0.82 and as low as 0.06 on the Thoughtful domain. The study concludes that current open-weight models cannot guarantee the necessary test constraint satisfaction for dependable automated model repair.

arXiv AI
Sep 11

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

The paper presents an end‑to‑end pipeline for translating natural language planning descriptions into PDDL problem instances using large language models. It incorporates multiple checks—syntactic parsing, planner success, domain conformance, an LLM critic, and iterative repair—to ensure faithfulness to the original task. Experiments on Planetarium, AutoPlanBench, and curated PDDL~2.1 problems reveal that operational success can diverge from benchmark‑reference reconstruction, and that structured repair improves outcomes while PDDL~2.1 remains challenging for reference reconstruction.

By Joana Rosa, Pedro Santos, Valdemar Oliveira, Rom\~ao Silva, L. Miguel Silveira, Bruno Martins
arXiv AI
Sep 17

Which LLM is Best for Translating Natural Language Goals to PDDL

The paper evaluates how well current Large Language Models can translate natural language goals, written by video game testers, into well‑formed PDDL targets for classical planning. Using a carefully designed prompt template, six state‑of‑the‑art LLMs were tested on correctness, speed, and error tendencies with real‑world benchmarks. All models achieved high correctness (>92%), with Gemini 2.5 Flash reaching 96% accuracy and the fewest false positives, while GPT‑4.1 was the fastest, yet differences in performance and occasional failures due to ambiguity and domain limits remain.

By Tomas Balyo, Lukas Chrpa, G. Michael Youngblood
arXiv AI
Sep 1

Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study

The paper presents a systematic analysis of five state‑of‑the‑art automated program repair agents, tracing their decision‑making across 500 real‑world repair tasks. It finds that while the agents perform well on simple fixes, they struggle with logic‑intensive bugs, often producing verbose, overfitted patches that pass tests without addressing root causes. Key bottlenecks identified include poor test generation, limited regression test selection, and reliance on primitive tooling without access to debuggers or advanced program analysis tools.

By Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
arXiv AI
6d ago

SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning

SLMFix is a code‑generation pipeline that uses a small language model fine‑tuned with reinforcement learning to correct syntactic errors in programs produced by large language models for domain‑specific languages. The approach relies on interpreter feedback to guide the error‑fixing process. Experiments show that SLMFix improves validator pass rates by 40% on low‑resource programming languages and removes over 50% of syntactic errors on high‑resource DSLs, outperforming supervised fine‑tuning even for 7B models.

By David Jiahao Fu, Aryan Gupta, Aaron Councilman, Yu-Xiong Wang, Vikram Adve
arXiv AI
Sep 24

Provably Complete Generalized Planning with LLMs

The paper presents a method for automatically generating generalized plans in Lean, along with formal proofs of their completeness for given domain constraints. It introduces a semantic‑preserving conversion from PDDL to Lean and uses an LLM to produce both the plan and its proof, whose correctness is verified by Lean’s kernel. Evaluated on 13 benchmark domains with GPT‑5.6‑Sol, the approach yields complete plans and valid proofs for 12 of them, marking a significant advance in automated generalized‑plan completeness.

By Katharina Stein, Chaahat Jain, J\"org Hoffmann, Alexander Koller
Hugging Face Trending Papers
Sep 24

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how Large Language Models (LLMs) handle bug fixing compared to human-written patches by analyzing about 3,000 Codeforces submissions. It finds that LLMs often modify more lines than necessary and sometimes produce entirely new solutions, and that they solve more problems correctly when generating solutions from scratch rather than patching existing code. The study highlights implications for AI‑assisted programming tools, suggesting a shift toward incremental problem‑solving strategies.

arXiv AI
Sep 18

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.

By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic