arXiv:2607. 10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments.
By Yixiong Chen, Xinyi Bai, Alan Yuille
arXiv:2606. 24162v1 Announce Type: cross Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics.
By Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei
The paper proposes a metric framework to differentiate cognitive amplification—where AI enhances human performance without eroding human capability—from cognitive delegation, which relies heavily on AI reasoning. It introduces four metrics (CAI*, D, HRI, HCDR) and tests them in NetLogo simulations across various reliance and dependency scenarios. The results show that positive collaborative gain is only achievable when an explicit interaction term is added, indicating that mere prevention of capability erosion is insufficient for genuine amplification.
By Eduardo Di Santi, Carla Florida
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Workin...
arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.
By Peter Tisnikar, Maja Swieczkowska, Benteng Ma, Gerard Canal, Matteo Leonetti
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2609.09134v1 Announce Type: new
Abstract: Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic...
By Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur
The paper introduces Recuris, a recursive Experiential‑Working Memory architecture that lets long‑horizon agents track task progress and select skills based on current needs rather than full history. By coupling working memory with experiential memory, execution becomes structured evidence that localizes failures to specific memory components, enabling a bounded recursive memory‑evolution loop. Across four benchmarks and ten models, Recuris improves task success in 35 of 37 model‑benchmark pairs, raising state‑of‑the‑art performance on tau‑bench and SkillFlow and reducing common long‑horizon failures by up to 80%.
By Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
CHIME introduces a credit‑aware hierarchical memory evolution framework that separates planning and execution experiences into distinct memory banks. By attributing each task outcome to the plan, execution, both, or neither before memorization, CHIME mitigates bias from noisy final outcomes and improves long‑horizon agent planning. Experiments on four benchmarks demonstrate that CHIME outperforms existing training‑based and self‑evolving memory methods, requires fewer memory items, and transfers effectively across backbone models.
By Yongshi Ye, Tian Lan, Feihu Jiang, Muyang Ye, Bin Zhu, Qianghuai Jia, Longyue Wang, Zhao Xu, Weihua Luo, Xiaodong Shi
arXiv:2604. 09670v2 Announce Type: replace-cross Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals.
By Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
arXiv:2606. 23195v2 Announce Type: replace Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence.
By Zewen Liu
The paper proposes a hierarchical architecture for long-horizon language‑model agents that must operate over days or weeks without forgetting. It introduces three key components: time‑scale levels that store bounded summaries, a clocked tick as the basic action unit, and cascaded intelligence that escalates tasks to more capable models only after review failures. A ten‑day experiment demonstrated that the agent maintained continuity across context resets, adapted its behavior based on early knowledge, and identified where learned components could be integrated.
By Erik Nijkamp, Anurag Koul, Egor Pakhomov, Bo Pang