arXiv AI

AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment

arXiv AI
Aug 14

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv:2608. 12841v1 Announce Type: cross Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations.

By Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang
arXiv AI
Aug 12

Recovering Wasted Compute in Autoresearch Agents

arXiv:2608. 10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch.

By Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao, Zaiqian Chen, Kazem Meidani, C. Bayan Bruss, Micah Goldblum
arXiv AI
Sep 25

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

AlphaDiverse is a framework that enhances large language model–based multi‑agent systems for alpha factor mining by addressing cost, availability, and confidentiality constraints. It generates diverse research paths through varied environments and post‑training local agents, then fine‑tunes these agents with supervised learning and optimizes them jointly using a GRPO method that balances predictive quality and diversity. The approach limits research feedback to inner‑period data and evaluates a frozen model on outer‑period data to avoid test‑set tuning, demonstrating competitive prediction and broader exploration across four Chinese stock universes.

By Qingzhuo Wang, Zikun Wei, Zhihua Wei, Wen Shen
arXiv Machine Learning
5d ago

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

The paper reports on applying AutoResearch—a large language model that iteratively edits training scripts—to optimize embedding systems for a book recommendation pipeline at production scale. Over twelve weeks, the authors ran 220+ experiments across two representation‑learning systems, uncovering five recurring failure modes (infrastructure fragility, agent memory decay, search‑direction stagnation, iteration‑cost asymmetry, and metric fixation) that were not present in smaller settings. They propose a three‑principle scaffolding (prevent, persist, redirect) to address these failures, achieving a 1.82× lift in Recall@6, a 2.1× lift in coherence, and an autonomous text‑only fallback that expanded catalog coverage by 5.8×.

By Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
arXiv AI
Aug 25

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

The article argues that agentic auto‑research should be guided by dense, intermediate signals of epistemic progress rather than by sparse final benchmarks. It compares this approach to fuzz testing, where coverage provides continuous feedback that directs input mutation. The authors propose controlled experiments to test whether such signals improve discovery efficiency and reduce false positives, and demonstrate in a simulated physics setting that an AI agent using feedback‑driven search uncovers a hidden law while optimization‑driven baselines fail.

By Yifeng He, Jicheng Wang, Yinzhe Zhao, Chengyang Shi, Jiachen Liu, Hao Chen
arXiv AI
Jun 4

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

arXiv:2606. 05080v1 Announce Type: new Abstract: Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts.

By Zhangchen Xu, Junda Chen, Yue Huang, Dongfu Jiang, Jiefeng Chen, Hang Hua, Zijian Wu, Zheyuan Liu, Zexue He, Lichi Li, Shizhe Diao, Jiaxin Pei, Jinsung Yoon, Hao Zhang, Mengdi Wang, Radha Poovendran, Misha Sra, Alex Pentland, Zichen Chen
arXiv AI
Sep 10

AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

AutoFyn is a new agent framework based on the Expert Iteration algorithm that updates a frozen model’s persistent state using verified reward signals instead of changing model weights. Each iteration starts with a fresh model session, and only durable information—such as memory files, reports, and repository state—is carried over through explicit interfaces. The system orchestrates exploration, planning, and specialized agents, while a task‑grounded verifier supplies objective rewards that are distilled back into the persistent state to guide the next round. AutoFyn has been applied to olympiad mathematics, data science, and cybersecurity, outperforming baseline models on the 2026 International Mathematical Olympiad, topping the Spider 2.0 dbt benchmark, and generating sixteen maintainer‑confirmed vulnerability advisories across several popular software projects.

By Adib Hasan, Daniel Schaffield, Akashnil Dutta, Tarik Adnan Moon