arXiv AI

Efficiency Matters in Autonomous Research

arXiv:2607. 24647v1 Announce Type: new Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks.

arXiv Machine Learning
Sep 23

Recursive self-improvement of AI research agents

The paper introduces AIDE^2, an AI research agent that recursively improves its own code by proposing, benchmarking, and selecting modifications. Over an eight‑day autonomous run, it achieved seven successive improvements—including new search policies and memory mechanisms—that transferred to four held‑out benchmarks in machine learning, algorithm engineering, and weather forecasting. The agent’s best version matched or outperformed a top human‑engineered production research agent and also reduced reward‑hacking rates, despite never optimizing for that metric.

By Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang
arXiv AI
Aug 17

AI Research Preference Models

arXiv:2608. 13940v1 Announce Type: new Abstract: AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days of GPU time.

By Thomas Simon Foster, Bassel Al Omari, Tingchen Fu, Thomas Mann, Carl Domond, Lucia Cipolina-Kun, Bhavul Gauri, Muna Aghamelu, Alexander D. Goldie, Eryk Helenowski, Jean-Christophe Gagnon-Audet, Alberto Pepe, Saba Nazir, Daniel Izcovich, Noam Levi, Rishi Hazra, Karen Hambardzumyan, Nicolas Baldwin, Xian Li, Martin Josifoski, Paris Giampouras, Masoud Jalili Sabet, Anya Sims, Hela Momand, Tatiana Shavrina, Despoina Magka, Jason Weston, Yulin Wang, Anirudh Goyal, Jo\~ao Henriques, Yoram Bachrach, Emily McMilin, Jakob Nicolaus Foerster
arXiv AI
Sep 17

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

PrimeScientist is a system that jointly selects research directions and allocates resources for autonomous research agents. It models the problem as a sequential decision task, using an executable plan tree to track competing plans and an adaptive MCTS-based policy to balance exploration and exploitation based on remaining resources and experimental feedback. Experiments on AI research, systems, code optimization, and machine learning engineering show that PrimeScientist improves average reward by 10.3% while reducing research attempts by 50.6% compared to AutoResearch under the same budget.

By Xinle Yu, Fan Bai, Kaiser Sun, Hengshuo Miao, Abhay Anand, Zhongyan Luo, Kun Zhou, Zhen Wang
arXiv AI
Aug 26

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

The paper introduces ExTS, a tree‑search policy designed for budget‑constrained agentic search where evaluation and generation costs are high. ExTS treats expansion as a value‑of‑information decision, combining discriminative reward shaping, a stochastic virtual child, and quality‑conditioned branching to allocate budget more effectively. Experiments on prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization show ExTS matching or surpassing task‑specific baselines with an average gain of +5.5% using a single configuration, and the authors also present pilot‑run diagnostics to guide adaptation to different problem structures.

By Haoyang Fang, Bernie Wang
arXiv AI
Jun 11

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

arXiv:2606. 11926v1 Announce Type: cross Abstract: Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction.

By Jiajie Jin, Yuyang Hu, Kai Qiu, Qi Dai, Chong Luo, Guanting Dong, Xiaoxi Li, Tong Zhao, Xiaolong Ma, Gongrui Zhang, Zhirong Wu, Bei Liu, Zhengyuan Yang, Linjie Li, Lijuan Wang, Hongjin Qian, Yutao Zhu, Zhicheng Dou
arXiv AI
Sep 18

Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

The paper critiques the common practice of evaluating large‑language‑model (LLM) evolutionary search methods using a single seed and fixed iteration budget, arguing that this approach is insufficient. By testing three search strategies across five optimization tasks and varying both the number of seeds (width) and iterations (depth), the authors find that optimal budget allocation depends on the strategy, task, and total budget, and that strategy rankings shift with different budgets. They propose a measurement protocol that maps the seeds‑by‑iterations frontier and offers practical guidance for researchers.

By Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay
arXiv AI
Aug 17

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

arXiv:2608. 14354v1 Announce Type: new Abstract: Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration decisions, and computational resources.

By Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang, Shuai Mao, Ting Lingya, Linyi Zou, Tailin Zhou, Yun Hin Chan, Wenkai Zhang, Zhanhong Zhou, Guowei Huang, Hongliang Li, Wenjing Cun, Zhitang Chen, Mingxuan Yuan, Yanhui Geng