arXiv:2606. 03544v1 Announce Type: new Abstract: Self-improving language agents are typically evaluated in isolation: an agent attempts a task, receives feedback, and iteratively refines its own behavior.
By Linyue Pan, Yaoming Zhu, Lin Qiu, Xuezhi Cao, Xunliang Cai
arXiv:2608. 09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents.
By Keyu He, Xuhui Zhou, Maarten Sap
arXiv:2607. 14574v1 Announce Type: new Abstract: Collective problem solving often requires that group members consider the tradeoff between exploitation of known solutions and exploration for new ones, where information of known solutions can be disseminated among individual members through communication networks.
By Hao He, Chris J. Kuhlman, Xinwei Deng
arXiv:2609. 37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures.
By Jungwoo Yang, In Jin Kong, Yohan Jo
arXiv:2607. 05297v1 Announce Type: new Abstract: Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability.
By Zefeng Wang, Minxi Yan, Jinhe Bi, Sikuan Yan, Volker Tresp, Yunpu Ma
CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.
By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
Rep2Skill introduces a representation-guided framework that enables large language model agents to self-evolve their textual skills by analyzing internal representation trajectories from agent rollouts. The method identifies execution turns that deviate from successful dynamics and uses these signals, together with execution contexts, as actionable feedback for targeted skill revision. Experiments with two open-source LLMs across two agent environments demonstrate that Rep2Skill consistently outperforms purely text-based approaches, showing that incorporating internal representations can enhance agent self-improvement.
By Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
arXiv:2609.38334v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-mode...
By Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng
The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.
By Mehmet Kerem Turkcan
arXiv:2604. 07821v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation fails.
By Advait Yadav, Sid Black, Oliver Sourbut
AREX-2 is a new approach that enhances the self‑improving ability of large language model agents by combining reflection—producing better solutions—and long‑horizon execution—maintaining effectiveness over many iterations. The method synthesizes improvement trajectories from machine‑learning and algorithmic programming tasks, providing verifiable feedback and sustained iteration. Trained on this data, an agent based on Qwen3.8‑27B achieves strong performance across multiple benchmarks and continues to improve as more iterative rounds are allowed.
By Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu
arXiv:2609.36675v1 Announce Type: new
Abstract: Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize extern...
By Ziqi Zhao, Fanqing Meng, Haocheng Lu, Lingxiao Du, Qiguang Chen, Mengkang Hu, Xiao-Ming Wu