arXiv AI By Hassan Sartaj, Shaukat Ali, Paolo Arcaini, Andrea Arcuri

Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap

Read the original on arXiv AI →

The paper reviews the state of search‑based software engineering (SBSE) and its relationship with emerging AI foundation models (FMs) such as large language models. It outlines a research roadmap that examines how FMs can enhance SBSE, how SBSE can advance FMs, and how the two can be integrated. The authors also propose future research directions and opportunities for applying SBSE in new domains enabled by FMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 8

Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

arXiv:2607. 06229v1 Announce Type: cross Abstract: Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filtering, sentiment analysis, extraction, similarity search, and aggregation within ordinary SQL queries.

By Tianyang Liu, Canwen Xu, Fangyu Lei, Nikki Lijing Kuang, Jixuan Chen, Tao Yu, Julian McAuley, Zhewei Yao, Yuxiong He
arXiv AI
Jun 2

SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned

arXiv:2602. 07666v4 Announce Type: replace-cross Abstract: DARPA's AI Cyber Challenge (AIxCC, 2023--2025) is the largest competition to date for building fully autonomous cyber reasoning systems (CRSs) that leverage recent advances in AI -- particularly large language models (LLMs) -- to discover and remediate vulnerabilities in real-world open-source software.

By Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Michael Pelican, David J. Musliner, Jeff Huang, Jon Silliman, Mikel Mcdaniel, Jefferson Casavant, Isaac Goldthwaite, Nicholas Vidovich, Matthew Lehman, Taesoo Kim
arXiv AI
5d ago

Learning from Research: Toward Lifelong Agent Harness Evolution

The paper introduces ScholarEvolve, a framework that evolves the software harness of language agents by automatically incorporating insights from recent research papers. It organizes harness improvements into functional modules, uses topic modeling to identify distinct strategies, and evaluates combinations to boost task performance. Experiments show significant gains on AppWorld and Tau2-Bench, raising Qwen3.5-27B completion rates from 49.6% to 63.6% and GPT-5.4-mini pass@1 from 72.7% to 81.9%.

By Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang
arXiv Machine Learning
Jun 24

Sakana Fugu Technical Report

arXiv:2606. 21228v2 Announce Type: replace Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains.

By Yujin Tang, Edoardo Cetin, Jinglue Xu, Qi Sun, Stefan Nielsen, Vincent Richard, Haruto Goda, Iaroslav Tymchenko, Nhan Nguyen, Hyunin Lee, Mari Ashiga, Shashank Kotyan, So Kuroki, Tarin Clanuwat