arXiv Computation and Language By Zihao Deng, Yining Zhu, Leiming Wang, Junbo Wang, Jingfei Lu

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

Read the original on arXiv Computation and Language →

Tree-of-Experience (ToE) is a hierarchical experience-management framework designed for large language model agents that aligns stored experiences with the agents’ reasoning hierarchy. By organizing experiences into a shared tree of analytical perspectives and reasoning paths, ToE calibrates reliability through environmental outcomes, enabling systematic updating, cross-task transfer, and efficient retrieval. Experiments on Game of 24 and FinEvolveBench demonstrate that ToE yields significant performance gains—31.4% accuracy improvement on Game of 24 and a 41.24% average improvement in tsIC on FinEvolveBench—outperforming both experience-free baselines and conventional experience-management methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 7

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.

By Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li
arXiv AI
Aug 20

SPADE: Self-Play in Adaptive Synthetic Executable Environments

SPADE (Self-Play in Adaptive Synthetic Executable Environments) is a reinforcement‑learning framework where a single large language model acts as both an Environment Designer—creating executable, long‑horizon training environments—and a Reasoning Agent—learning to act within those environments. The framework uses a regret signal based on the difference between rewarded performance with and without privileged hints to guide the Designer toward environments that are challenging yet solvable. Experiments show that, when scaled to 30‑billion‑parameter models, SPADE outperforms fixed‑environment baselines by significant margins across math, science, code, and reasoning benchmarks, and improves tool‑use performance on BFCL‑v4 and ACEBench‑Agent. whyItMatters":"By making environment design a learnable component, SPADE enables continuous self‑improvement and demonstrates that adaptive, self‑generated training environments can substantially boost language‑model performance across diverse tasks."

By Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques