Tree-of-Experience (ToE) is a hierarchical experience-management framework designed for large language model agents that aligns stored experiences with the agents’ reasoning hierarchy. By organizing experiences into a shared tree of analytical perspectives and reasoning paths, ToE calibrates reliability through environmental outcomes, enabling systematic updating, cross-task transfer, and efficient retrieval. Experiments on Game of 24 and FinEvolveBench demonstrate that ToE yields significant performance gains—31.4% accuracy improvement on Game of 24 and a 41.24% average improvement in tsIC on FinEvolveBench—outperforming both experience-free baselines and conventional experience-management methods.
By Zihao Deng, Yining Zhu, Leiming Wang, Junbo Wang, Jingfei Lu
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve.
By Akshay Nambi, Yash Pandya, Sahil Gupta, Sarthak Harne, Kavyansh Chourasia, Yash Lara, Ahmed Awadallah, Ece Kamar
arXiv:2606. 04703v1 Announce Type: cross Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward continual learning in large language models (LLMs).
By Jingwen Chen, Wenkai Yang, Shengda Fan, Wenbo Nie, Chenxing Sun, Shaodong Zheng, Yangen Hu, Lu Pan, Ke Zeng, Yankai Lin
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deployment is inherently dynamic, requiring agents to continually align their knowledge, skills, and behavior with changing environments and updated task conditions.
arXiv:2608. 12428v1 Announce Type: new Abstract: Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions.
By Kaichao Liang, Yuqi Cui, Hao Kong, Xinyuan Huang, Guohaotian Hou, Qingcan Kang, Liang Chen, Yiyang Yin, Ke Ye, Jiaquan Guo, Da Chen, Lingan Zeng, Yixing Peng, Rong Yao, Shixiong Kai, Mingxuan Yuan
arXiv:2608. 15165v1 Announce Type: new Abstract: Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge.
By Yu He, Weikai Yang
arXiv:2608. 16544v1 Announce Type: cross Abstract: Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules.
By Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
arXiv:2608. 15451v1 Announce Type: new Abstract: Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations.
By Oliver Kramer
The paper introduces an online skill‑evolution framework that transforms interaction traces and evaluator feedback into a persistent, versioned library of reusable procedures for computer‑use agents. By executing each iteration against a frozen library snapshot, the system updates skills without altering the underlying model parameters. Experiments across four OSWorld domains show that the evolving library consistently outperforms an empty‑library baseline, with gains ranging from 5.7 to 18.6 percentage points, while also revealing domain‑specific temporal stability and challenges in skill retrieval and revision.
By Longtao Hu, Xiao Liang, Linchao Zhu
arXiv:2606. 07603v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through task interactions.
By Bowen Ren, Heyan Huang, Yinghao Li, Yang Gao
arXiv:2603. 20667v2 Announce Type: replace-cross Abstract: Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks.
By Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan, Arman Cohan
arXiv:2605. 27366v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills as isolated, static artifacts, limiting reusability, reliability, and long-term improvement.
By Huawei Lin, Peng Li, Jie Song, Fuxin Jiang, Tieying Zhang