arXiv AI

AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems

arXiv:2606. 15834v1 Announce Type: new Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems.

arXiv AI
Sep 17

A Study of the Reliability of Agentic AI-Generated Programs

The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.

By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
arXiv AI
Jun 8

EvoClaw: Evaluating AI Agents on Continuous Software Evolution

arXiv:2603. 13428v2 Announce Type: replace-cross Abstract: With AI agents increasingly deployed as long-running systems, it becomes essential to autonomously construct and continuously evolve customized software to enable interaction within dynamic environments.

By Gangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan, Yuhong Liu, Yuxin Yang, Dhruv Parikh, Rajgopal Kannan, Le Cong, Mengdi Wang, Qian Zhang, Viktor Prasanna, Xiangru Tang, Xingyao Wang
arXiv AI
6d ago

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

HARDEN is a constrained evolutionary search method that transforms existing evaluation cases into more challenging variants while preserving their expected outputs. It operates along domain‑specific complexity axes and enforces feasibility constraints such as task semantics, realism, and execution validity. Experiments on FinQA, PubMedQA, and ContractNLI with Qwen3.5 models show that HARDEN can reduce task‑model accuracy by an average of 22.7% and up to 49.9% compared to single‑pass baselines.

By Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar
arXiv AI
Jun 2

MemPro: Agentic Memory Systems as Evolvable Programs

arXiv:2606. 00619v1 Announce Type: cross Abstract: Long-horizon autonomous agents require memory systems to retain historical information, track evolving states, and reuse relevant knowledge beyond finite context windows.

By Qingshan Liu, Guoqing Wang, Wen Wu, Jingqi Huang, Xinqi Tao, Dejia Song, Jie Zhou, Liang He
arXiv Machine Learning
1d ago

TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

TRACE tackles real‑world dynamic resource assignment by combining evolutionary automatic heuristic design with an agentic knowledge‑extraction workflow. A Reasoner agent interprets system logs to hypothesize about underlying dynamics, while a Coder agent generates and runs schema‑specific code to validate these hypotheses, producing insights or executable tools for the evolved heuristics. Evaluations on a synthetic cloud benchmark and a 5G vRAN scenario show that TRACE outperforms existing AHD methods, delivering more auditable heuristics with less than 2% overhead.

By Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez