Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Workin...
ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.
By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren
arXiv:2606. 02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
arXiv:2606. 02461v2 Announce Type: replace Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
By Yiheng Shu, Bernal Jim\'enez Guti\'errez, Saisri Padmaja Jonnalagedda, Yuguang Yao, Huan Sun, Yu Su
The paper introduces Recuris, a recursive Experiential‑Working Memory architecture that lets long‑horizon agents track task progress and select skills based on current needs rather than full history. By coupling working memory with experiential memory, execution becomes structured evidence that localizes failures to specific memory components, enabling a bounded recursive memory‑evolution loop. Across four benchmarks and ten models, Recuris improves task success in 35 of 37 model‑benchmark pairs, raising state‑of‑the‑art performance on tau‑bench and SkillFlow and reducing common long‑horizon failures by up to 80%.
By Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
By Anne Harrington, Nayan Saxena, Michael Murphy, Anastasia Borovykh, Zeyu Yun, Sridhar Kamath, Ara Eindra Kyi, Trevor Darrell, Jitendra Malik, Yutong Bai
The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the underlying foundation model frozen. HCL defines harness-level forgetting and proposes a guarded evolution process involving a Continual Optimizer and Evaluator to ensure improvements without losing prior behavior. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate how the stability–plasticity trade‑off can be explicitly tuned.
By Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao
arXiv:2608. 03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks.
By Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
Long‑horizon language model agents accumulate reasoning history, which inflates context length and inference cost. The paper introduces Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training‑free online method that ranks and removes reasoning blocks based on frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR raises average reward from 0.699 to 0.718 and cuts input, output, and cache read tokens by 25.5%, 14.4%, and 33.3% respectively, while analyses show that historical reasoning becomes replaceable once task‑relevant state is externalized.
By Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the foundation model frozen. HCL defines harness-level forgetting and proposes guarded harness evolution with a Continual Optimizer and Evaluator to balance improvement, retention, and validity. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate explicit control over the stability–plasticity trade‑off.
arXiv:2609.17416v1 Announce Type: new
Abstract: Voice agents built on LLMs follow a rigid listen-think-speak loop that inserts seconds of dead air before every reply. We show that continuous-time cog...
By Bojie Li, Noah Shi
arXiv:2609.05435v2 Announce Type: replace
Abstract: Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability thr...
By Zerui Cheng, Jiawei Xu, Huacan Chai, Jiayang Sun, Pramod Viswanath, Maxm Pan