Hugging Face Trending Papers

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions.

arXiv Machine Learning
Sep 10

MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves

arXiv:2609.06396v2 Announce Type: new Abstract: Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. Yet...

By Zihan Tan, Leixin Sun, Zitong Shi, Yitao Liu, Jiajun Wu, Nathaniel Brooks, Jiaru Qian, Xiaoran Shang, Suyuan Huang, Yi Ding, Yangxu Liao, Mukai Li, Qiushi Sun, Shudong Liu, Xuankun Rong, Xiaohang Yu, Zhuo Chen, Hejia Geng, Chenxin Li, Aozhou Wang, Zengji Tu, Robert Tang, Yuxin Zhan, Eric Jiang, Yuxin Wu, Jianqing Zhang, Xiao Liang, Fang Wu, Haochi Zhang, Alexander Marlow, Guancheng Wan
arXiv Machine Learning
Sep 23

Recursive self-improvement of AI research agents

The paper introduces AIDE^2, an AI research agent that recursively improves its own code by proposing, benchmarking, and selecting modifications. Over an eight‑day autonomous run, it achieved seven successive improvements—including new search policies and memory mechanisms—that transferred to four held‑out benchmarks in machine learning, algorithm engineering, and weather forecasting. The agent’s best version matched or outperformed a top human‑engineered production research agent and also reduced reward‑hacking rates, despite never optimizing for that metric.

By Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang
arXiv AI
Sep 15

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

The paper introduces Generalized Agent Iteration (GAI), a formal framework that unifies iterative policy improvement and recursive self‑improvement (RSI) under a single learning paradigm. GAI treats an agent as a configuration of modifiable components and models learning as a cycle of evaluation and improvement, with two key dials: whether the improving mechanism is part of the agent and whether the evaluation standard is external. These dials distinguish between generalized policy iteration (GPI) and RSI, and classify systems as anchored, goal‑drift, or fully self‑referential, allowing existing systems to be mapped and RSI defects to be analyzed systematically.

By Hongyao Tang, Yi Ma, Pengyi Li, Yifu Yuan
arXiv AI
Jul 20

Recursive Harness Self-Improvement

arXiv:2607. 15524v1 Announce Type: cross Abstract: Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models.

By Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang
arXiv Machine Learning
Sep 11

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

The paper discusses recursive self‑improvement (RSI) for AI, describing how systems can use experience and feedback to make lasting enhancements to both their abilities and their future improvement processes. It introduces the Headroom‑Closed Index (HCI) to expose limitations in current large language models and outlines a development roadmap for RSI, progressing from autonomy in execution to full recursive meta‑improvement. The authors analyze RSI in various contexts such as scientific discovery, embodied intelligence, and software engineering, noting differing requirements and development speeds, and they connect RSI research to practical industry applications while highlighting key challenges to achieving genuine RSI.

By Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou, Yumou Liu, Bangrui Xu, Yukai Wu, Sidi Chen, Yuhan Zhou, Haoyu Wang, Xiaoyou Yu, Shaokun Han, Xuzhou Zhu, Le Zhou, Bolin Lu, Wei Zhou, Jiachen Liu, Nuozhou Fang, Jiaxin Tian, Ruoyu Chen, Yuxuan Li, Kai Zuo, Kaiyan Zhang, Jiantao Qiu, Conghui He, Guoliang Li, Bowen Zhou, Zhiyuan Liu, Zhoufutu Wen, Jihua Kang, Xuanhe Zhou, Fan Wu
arXiv AI
Sep 25

iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

The paper introduces iCoder-27B, a 27‑billion‑parameter model for RTL design and GPU kernel optimization that is developed through a recursive AI‑led process with minimal human input. Human experts provide high‑level objectives and reusable research skills, while the agent autonomously selects experiments, diagnoses outcomes, and refines training strategies, coordinating SFT, self‑distillation, and reinforcement learning. iCoder outperforms GPT‑5.5 and Claude‑Opus‑4.8 on several benchmarks, demonstrating the feasibility of building frontier‑competitive models with largely automated development.

By Cheng Yang, Jiayang Lyu, Shangyuan Liu, Guibin Zhang, Jiong Lin, Xinlei Yu, Junchi Yan, Shuicheng Yan, Weinan E, Linfeng Zhang, Linfeng Zhang, Qibing Ren