The paper demonstrates that large language models can generate executable procedural content generators, enabling direct search over generator programs rather than individual levels. Using Sokoban, Zelda, Dangerous Dave, and Lode Runner, the authors evolve complete Python generators via language‑model mutation and crossover, and introduce Continual Abstraction Discovery (CAD) to extract reusable primitives into a run‑specific helper module. Experiments show that CAD consistently improves mean final best fitness across all domain and API comparisons, with learned libraries being adopted by subsequent programs and repeatedly rediscovering useful utilities.
By Matthew Siper, Ahmed Khalifa, Julian Togelius
arXiv:2607. 00062v1 Announce Type: cross Abstract: High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorithms.
By Xinyuan Song, Zekun Cai, Liang Zhao
arXiv:2505. 03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics.
By Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks and domains. However, existing LLM-driven evolutionary frameworks largely discard such knowledge, repeatedly rediscovering similar ideas and limiting opportunities for cross-run and cross-task learning.
arXiv:2609.22878v1 Announce Type: new
Abstract: Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant tr...
By Aditya Pola, Arkaprava Majumdar, Vineeth N. Balasubramanian
arXiv:2606. 15834v1 Announce Type: new Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems.
By Yajie Zhou, Ao Li, Ashwin Silla, Zaoxing Liu, Vyas Sekar