arXiv AI

Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap

The paper reviews the state of search‑based software engineering (SBSE) and its relationship with emerging AI foundation models (FMs) such as large language models. It outlines a research roadmap that examines how FMs can enhance SBSE, how SBSE can advance FMs, and how the two can be integrated. The authors also propose future research directions and opportunities for applying SBSE in new domains enabled by FMs.

arXiv AI
Jul 8

Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

arXiv:2607. 06229v1 Announce Type: cross Abstract: Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filtering, sentiment analysis, extraction, similarity search, and aggregation within ordinary SQL queries.

By Tianyang Liu, Canwen Xu, Fangyu Lei, Nikki Lijing Kuang, Jixuan Chen, Tao Yu, Julian McAuley, Zhewei Yao, Yuxiong He
arXiv AI
Jun 2

SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned

arXiv:2602. 07666v4 Announce Type: replace-cross Abstract: DARPA's AI Cyber Challenge (AIxCC, 2023--2025) is the largest competition to date for building fully autonomous cyber reasoning systems (CRSs) that leverage recent advances in AI -- particularly large language models (LLMs) -- to discover and remediate vulnerabilities in real-world open-source software.

By Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Michael Pelican, David J. Musliner, Jeff Huang, Jon Silliman, Mikel Mcdaniel, Jefferson Casavant, Isaac Goldthwaite, Nicholas Vidovich, Matthew Lehman, Taesoo Kim
arXiv AI
5d ago

Learning from Research: Toward Lifelong Agent Harness Evolution

The paper introduces ScholarEvolve, a framework that evolves the software harness of language agents by automatically incorporating insights from recent research papers. It organizes harness improvements into functional modules, uses topic modeling to identify distinct strategies, and evaluates combinations to boost task performance. Experiments show significant gains on AppWorld and Tau2-Bench, raising Qwen3.5-27B completion rates from 49.6% to 63.6% and GPT-5.4-mini pass@1 from 72.7% to 81.9%.

By Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang
arXiv Machine Learning
Jun 24

Sakana Fugu Technical Report

arXiv:2606. 21228v2 Announce Type: replace Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains.

By Yujin Tang, Edoardo Cetin, Jinglue Xu, Qi Sun, Stefan Nielsen, Vincent Richard, Haruto Goda, Iaroslav Tymchenko, Nhan Nguyen, Hyunin Lee, Mari Ashiga, Shashank Kotyan, So Kuroki, Tarin Clanuwat
arXiv AI
Sep 7

AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

AutoLR is an autonomous harness designed to streamline the iterative research‑and‑engineering cycle for industrial recommender systems, exemplified by NetEase’s gaming‑community app DASHEN. It integrates a multi‑expert council for adversarial review, a deterministic evidence‑weighted selector to allocate trial budgets, and a layered knowledge system that fuses external research with domain‑specific insights and empirical evidence. Large language model agents handle semantic reasoning and code generation, while deterministic controllers maintain control over execution, metrics, guardrails, and state management.

By Qi Zhang, Yanlin Chen, Wenchao Xiao
arXiv AI
Sep 4

Evolving Excellence: Automated Optimization of LLM-based Agents

The paper introduces ARTEMIS, a no-code evolutionary optimization platform that automatically tunes large language model (LLM) agents by jointly optimizing prompts, tool descriptions, and parameters using semantically-aware genetic operators. Starting from a benchmark script and natural language goals, ARTEMIS discovers configurable components, extracts performance signals from execution logs, and evolves configurations without architectural changes. Experiments on four agent systems show significant gains: a 13.6% increase in acceptance rate for the ALE Agent, a 10.1% performance boost for the Mini‑SWE Agent, a 36.9% token‑reduction for the CrewAI Agent, and a 22% accuracy improvement for the MathTales‑Teacher Agent using a smaller open‑source model.

By Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, Tianhua Xu, Aurora Constantin, Leslie Kanthan, Michail Basios
arXiv AI
6d ago

Towards an AI Software Factory for Data Systems

The paper describes progress toward an AI Software Factory that accelerates every stage of the software development lifecycle—Targeting, Coding, Reviewing, and Ops—by generating a metadata exhaust for self‑improvement. It focuses on data systems and evolutionary coding tasks, reporting scaled deployments at Microsoft that achieved three‑fold engineering efficiency over agentic coding and up to 22‑fold token efficiency. The authors also outline several open challenges for further development.

By Anna Pavlenko, Bogdan Crivat, Brandon Haynes, Carlo Curino, Fotis Psallidas, Jaro Slawinski, Johannes Freischuetz, Laura Pereira Sanchez, Markus Weimer, Mathieu Demarne, Matthias Jasny, Mauktik Gandhi, Max Bovykin, Mirco Milletari, Purbasha Ghosh, Qiushi Bai, Raghu Ramakrishnan, Rahul Pandita, Sergiy Matusevich, Shivaram Venkataraman, Subru Krishnan, Md. Tareq Mahmood, Tiemo Bang, Venkatesh Emani, Xuan Zhao, Yiwen Zhu
arXiv AI
Jul 7

Gypscie: A Cross-Platform AI Artifact Management System

arXiv:2604. 10311v2 Announce Type: replace Abstract: Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more advanced approaches such as deep learning and large language models (LLMs), play a central role in modern applications.

By Fabio Porto, Eduardo Ogasawara, Gabriela Moraes Botaro, Julia Neumann Bastos, Augusto Fonseca, Esther Pacitti, Patrick Valduriez