arXiv AI By Anna Pavlenko, Bogdan Crivat, Brandon Haynes, Carlo Curino, Fotis Psallidas, Jaro Slawinski, Johannes Freischuetz, Laura Pereira Sanchez, Markus Weimer, Mathieu Demarne, Matthias Jasny, Mauktik Gandhi, Max Bovykin, Mirco Milletari, Purbasha Ghosh, Qiushi Bai, Raghu Ramakrishnan, Rahul Pandita, Sergiy Matusevich, Shivaram Venkataraman, Subru Krishnan, Md. Tareq Mahmood, Tiemo Bang, Venkatesh Emani, Xuan Zhao, Yiwen Zhu

Towards an AI Software Factory for Data Systems

Read the original on arXiv AI →

The paper describes progress toward an AI Software Factory that accelerates every stage of the software development lifecycle—Targeting, Coding, Reviewing, and Ops—by generating a metadata exhaust for self‑improvement. It focuses on data systems and evolutionary coding tasks, reporting scaled deployments at Microsoft that achieved three‑fold engineering efficiency over agentic coding and up to 22‑fold token efficiency. The authors also outline several open challenges for further development.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Evolving Excellence: Automated Optimization of LLM-based Agents

The paper introduces ARTEMIS, a no-code evolutionary optimization platform that automatically tunes large language model (LLM) agents by jointly optimizing prompts, tool descriptions, and parameters using semantically-aware genetic operators. Starting from a benchmark script and natural language goals, ARTEMIS discovers configurable components, extracts performance signals from execution logs, and evolves configurations without architectural changes. Experiments on four agent systems show significant gains: a 13.6% increase in acceptance rate for the ALE Agent, a 10.1% performance boost for the Mini‑SWE Agent, a 36.9% token‑reduction for the CrewAI Agent, and a 22% accuracy improvement for the MathTales‑Teacher Agent using a smaller open‑source model.

By Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, Tianhua Xu, Aurora Constantin, Leslie Kanthan, Michail Basios
arXiv AI
Sep 10

FrogNano: Training a 4B Coding Agent via Online Task Synthesis

FrogNano is a 4B coding agent trained exclusively with reinforcement learning on about 1,500 synthetic software engineering environments. Its training leverages an online task synthesis pipeline that generates tasks at the current agent’s learnability frontier, improving performance without distilling from larger models. The report details the methodology, evaluates the agent across diverse environments, and analyzes its effectiveness as a lightweight coding agent for minimal hardware.

By Minseon Kim, Zhengyan Shi, Emiliano Penaloza, Christopher Cui, Roger Creus Castanyer, Maryam Hashemzadeh, Isadora White, Jonathan Light, Jeonghye Kim, Matheus Pereira, Darya Moldavskaya, Chinmay Singh, Fabio Vera, Baolin Peng, Xingdi Yuan, Marc-Alexandre C\^ot\'e, Alessandro Sordoni