arXiv AI

Recovering Wasted Compute in Autoresearch Agents

arXiv:2608. 10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch.

arXiv Machine Learning
5d ago

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

The paper reports on applying AutoResearch—a large language model that iteratively edits training scripts—to optimize embedding systems for a book recommendation pipeline at production scale. Over twelve weeks, the authors ran 220+ experiments across two representation‑learning systems, uncovering five recurring failure modes (infrastructure fragility, agent memory decay, search‑direction stagnation, iteration‑cost asymmetry, and metric fixation) that were not present in smaller settings. They propose a three‑principle scaffolding (prevent, persist, redirect) to address these failures, achieving a 1.82× lift in Recall@6, a 2.1× lift in coherence, and an autonomous text‑only fallback that expanded catalog coverage by 5.8×.

By Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
arXiv Machine Learning
Sep 22

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

The paper introduces Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning of large language models cumulative. In the first stage, a multi-agent pipeline uses Monte Carlo Tree Search to explore training strategies while a Distillation Agent records task-specific insights and cross-task confidence scores into a structured repository. In the second stage, SAGE retrieves relevant experience from this repository to guide training on new tasks, achieving a 12.4‑percentage‑point improvement over a baseline pipeline without accumulated experience on nine unseen tasks.

By Haoran Zhao, Wei Du, Dingwen Yang, Jixuan Huang, Junlin Shang, Lingyong Fang, Ya Guo, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Jun 4

Can Generalist Agents Automate Data Curation?

arXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback.

By Feiyang Kang, Hanze Li, Adam Nguyen, Mahavir Dabas, Jiaqi W. Ma, Frederic Sala, Dawn Song, Ruoxi Jia
arXiv AI
Aug 7

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv:2608. 05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers.

By Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
arXiv Computer Vision
Aug 28

OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks

OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.

By Jing Wu, Wenjie Ai, Daphne Barretto, Yiye Chen, Qingyu Chen, Yuhang He, Pranit Chawla, Nicholas Gyd\'e, Yanan Jian, Vibhav Vineet
arXiv Machine Learning
Sep 23

Recursive self-improvement of AI research agents

The paper introduces AIDE^2, an AI research agent that recursively improves its own code by proposing, benchmarking, and selecting modifications. Over an eight‑day autonomous run, it achieved seven successive improvements—including new search policies and memory mechanisms—that transferred to four held‑out benchmarks in machine learning, algorithm engineering, and weather forecasting. The agent’s best version matched or outperformed a top human‑engineered production research agent and also reduced reward‑hacking rates, despite never optimizing for that metric.

By Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang
arXiv AI
Aug 25

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

The article argues that agentic auto‑research should be guided by dense, intermediate signals of epistemic progress rather than by sparse final benchmarks. It compares this approach to fuzz testing, where coverage provides continuous feedback that directs input mutation. The authors propose controlled experiments to test whether such signals improve discovery efficiency and reduce false positives, and demonstrate in a simulated physics setting that an AI agent using feedback‑driven search uncovers a hidden law while optimization‑driven baselines fail.

By Yifeng He, Jicheng Wang, Yinzhe Zhao, Chengyang Shi, Jiachen Liu, Hao Chen