arXiv:2606. 02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
arXiv:2606. 02461v2 Announce Type: replace Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.
The paper proposes a hierarchical architecture for long-horizon language‑model agents that must operate over days or weeks without forgetting. It introduces three key components: time‑scale levels that store bounded summaries, a clocked tick as the basic action unit, and cascaded intelligence that escalates tasks to more capable models only after review failures. A ten‑day experiment demonstrated that the agent maintained continuity across context resets, adapted its behavior based on early knowledge, and identified where learned components could be integrated.
By Erik Nijkamp, Anurag Koul, Egor Pakhomov, Bo Pang
arXiv:2608. 10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants.
By Xiaohongshu Inc
arXiv:2609.17416v1 Announce Type: new
Abstract: Voice agents built on LLMs follow a rigid listen-think-speak loop that inserts seconds of dead air before every reply. We show that continuous-time cog...
By Bojie Li, Noah Shi
arXiv:2607. 03441v1 Announce Type: cross Abstract: LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked.
By Yanbo Wang, Jinhua Hao, Yuze Shi, Kun Yuan, Ming Sun
arXiv:2606. 05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions.
By Yunxiang Zhang, Yiheng Li, Ali Payani, Lu Wang
arXiv:2608.30478v1 Announce Type: new
Abstract: Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and decision-making procedures, enabling...
By Shihan Dou, Haoxiang Jia, Shichun Liu, Feng Chen, Chenhao Huang, Yujiong Shen, Shaofan Liu, Jiayi Chen, Jiahang Lin, Honglin Guo, Qianyu He, Minghao Guo, Ziyi Ye, Pluto Zhou, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2604. 19775v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments.
By Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy
The paper introduces a new multi‑agent micro‑benchmark called Delay‑of‑Gratification, modeled after the Stanford marshmallow experiment, to evaluate large language models (LLMs) in long‑horizon, multi‑turn interactions. In the benchmark, ReAct agents use a per‑step “raise a question” tool under various constraints—social context (broadcast vs. isolated), persona traits (age, hedonic drive), and tool‑use policy (mandatory vs. optional). Across 19,200 trajectories, the study finds that most agents exhibit an early impulse to “eat,” only 75.9% persist to the end, and factors such as isolation and hedonic drive significantly influence survival and questioning behavior, with ablations showing that removing hedonic drive and age can improve completion rates.
By Olga Manakina, Igor Bogdanov, Chung-Horng Lung
arXiv:2608. 00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience.
By Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan