arXiv:2607. 19592v1 Announce Type: new Abstract: Self-improving AI systems typically treat the agent as the object that improves, by optimizing prompts, workflows, harnesses, or even the agent's own code.
By Xuefei Julie Wang, Lauren Hyoseo Yoon, Chengrui Qu, Amanda Zichang Wang, Atharva Sehgal, Eric Mazumdar, Yisong Yue
SPADE (Self-Play in Adaptive Synthetic Executable Environments) is a reinforcement‑learning framework where a single large language model acts as both an Environment Designer—creating executable, long‑horizon training environments—and a Reasoning Agent—learning to act within those environments. The framework uses a regret signal based on the difference between rewarded performance with and without privileged hints to guide the Designer toward environments that are challenging yet solvable. Experiments show that, when scaled to 30‑billion‑parameter models, SPADE outperforms fixed‑environment baselines by significant margins across math, science, code, and reasoning benchmarks, and improves tool‑use performance on BFCL‑v4 and ACEBench‑Agent.
whyItMatters":"By making environment design a learnable component, SPADE enables continuous self‑improvement and demonstrates that adaptive, self‑generated training environments can substantially boost language‑model performance across diverse tasks."
By Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
The paper introduces ANCHOR, an external supervisory framework driven by large language models (LLMs) that provides evaluative feedback at multiple stages of self‑evolving agents. By integrating ANCHOR into two open‑source self‑evolving agent frameworks, the authors demonstrate that it significantly improves safety performance while preserving core capabilities across coding, mathematical reasoning, and safety tasks. The study also finds that supervision based on execution results is especially effective and that increasing supervision frequency yields diminishing returns, offering practical guidance for future research.
By Dianxing Shi, Bowen Wang, Junqi He, Junhao Chen, Yuta Nakashima
arXiv:2603. 20667v2 Announce Type: replace-cross Abstract: Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks.
By Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan, Arman Cohan
arXiv:2602. 07883v3 Announce Type: replace Abstract: LLM-powered agentic systems excel at complex long-horizon tasks, but remain constrained by static configurations fixed before execution.
By Jingqi Zhou, Sheng Wang, Dezhao Deng, Junwen Lu, Junwei Su, Qintong Li, Jiahui Gao, Hao Wu, Jiyue Jiang, Lingpeng Kong, Dunhong Jin, Chuan Wu
The paper introduces Generalized Agent Iteration (GAI), a formal framework that unifies iterative policy improvement and recursive self‑improvement (RSI) under a single learning paradigm. GAI treats an agent as a configuration of modifiable components and models learning as a cycle of evaluation and improvement, with two key dials: whether the improving mechanism is part of the agent and whether the evaluation standard is external. These dials distinguish between generalized policy iteration (GPI) and RSI, and classify systems as anchored, goal‑drift, or fully self‑referential, allowing existing systems to be mapped and RSI defects to be analyzed systematically.
By Hongyao Tang, Yi Ma, Pengyi Li, Yifu Yuan