The paper demonstrates that large language models can generate executable procedural content generators, enabling direct search over generator programs rather than individual levels. Using Sokoban, Zelda, Dangerous Dave, and Lode Runner, the authors evolve complete Python generators via language‑model mutation and crossover, and introduce Continual Abstraction Discovery (CAD) to extract reusable primitives into a run‑specific helper module. Experiments show that CAD consistently improves mean final best fitness across all domain and API comparisons, with learned libraries being adopted by subsequent programs and repeatedly rediscovering useful utilities.
By Matthew Siper, Ahmed Khalifa, Julian Togelius
arXiv:2607. 00062v1 Announce Type: cross Abstract: High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorithms.
By Xinyuan Song, Zekun Cai, Liang Zhao
arXiv:2505. 03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics.
By Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks and domains. However, existing LLM-driven evolutionary frameworks largely discard such knowledge, repeatedly rediscovering similar ideas and limiting opportunities for cross-run and cross-task learning.
arXiv:2609.22878v1 Announce Type: new
Abstract: Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant tr...
By Aditya Pola, Arkaprava Majumdar, Vineeth N. Balasubramanian
arXiv:2606. 15834v1 Announce Type: new Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems.
By Yajie Zhou, Ao Li, Ashwin Silla, Zaoxing Liu, Vyas Sekar
arXiv:2606. 26173v1 Announce Type: new Abstract: Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs.
By Dhruv Sharma, Gautam Shroff
AlgoEvo introduces a unified agentic framework for automated algorithm discovery that replaces rigid search pipelines with an interactive, knowledge‑accumulating process. An autonomous agent inspects, diagnoses, and edits code using runtime feedback, while a design skill hub decouples paradigm‑specific knowledge from the core engine, enabling a single workflow to handle single‑objective, multi‑objective, and multi‑component design tasks. The hierarchical experience mechanism organizes search trajectories into a task‑level tree, guiding exploration and consolidating cross‑task patterns into reusable skills, resulting in performance that matches or surpasses specialized methods with fewer evaluations and reduced token consumption.
By Junhao Qiu, Qinglong Hu, Xialiang Tong, Mingxuan Yuan, Liyong Lin, Qingfu Zhang
The paper introduces ARTEMIS, a no-code evolutionary optimization platform that automatically tunes large language model (LLM) agents by jointly optimizing prompts, tool descriptions, and parameters using semantically-aware genetic operators. Starting from a benchmark script and natural language goals, ARTEMIS discovers configurable components, extracts performance signals from execution logs, and evolves configurations without architectural changes. Experiments on four agent systems show significant gains: a 13.6% increase in acceptance rate for the ALE Agent, a 10.1% performance boost for the Mini‑SWE Agent, a 36.9% token‑reduction for the CrewAI Agent, and a 22% accuracy improvement for the MathTales‑Teacher Agent using a smaller open‑source model.
By Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, Tianhua Xu, Aurora Constantin, Leslie Kanthan, Michail Basios
arXiv:2507. 22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation.
By Qiushi Sun, Jinyang Gong, Lei Li, Qipeng Guo, Fei Yuan
arXiv:2605. 29649v2 Announce Type: replace Abstract: Heuristic search is the dominant paradigm in symbolic AI planning, and the strongest heuristics are the result of decades of work by planning researchers.
By Elliot Gestrin, Jendrik Seipp
EngramBench is a new benchmark designed to evaluate skill evolution in autonomous agents by focusing on genuine capability abstraction rather than solution copying. It includes 30 learning tasks and 13 unseen transfer tasks that require agents to manage complex, multi-hour development cycles with LLM‑simulated users. The study shows that while static skill banks cannot eliminate the need for precise code implementation, they effectively reduce redundant context and cut overall coding time by more than 55%.
By Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao