arXiv:2606. 10209v1 Announce Type: new Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can cause context overflow, stale-state errors, and high inference cost.
By Abhilasha Lodha, Mahsa Pahlavikhah Varnosfaderani, Abir Chakraborty, Abhinav Mithal
arXiv:2608. 04830v1 Announce Type: new Abstract: Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.
By Bo Wang, Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu, Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang, Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu
The paper evaluates five context‑trimming strategies for agentic large language model workflows, comparing them on metrics such as task success, protocol adherence, token savings, and latency. Conventional trimming methods save about 60% of tokens but achieve lower success rates, while protocol‑aware trimming raises success to 92.2% and adaptive guardrails further improve it to 96% success with 56% token savings. The study shows that preserving protocol‑critical state is more important than aggressive token removal, and that adaptive guardrails enhance efficiency, scalability, and reliability for long‑horizon agentic systems.
By Harish Gaggar
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams.
arXiv:2606. 18191v1 Announce Type: new Abstract: Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries.
By Md Tawkat Islam Khondaker, Raymond Li, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan, Issam H. Laradji
arXiv:2607. 26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings.
By Changyu Du, Alexander Vosseler, Filippo Mazza, Andr\'e Borrmann
arXiv:2606. 06337v1 Announce Type: new Abstract: Large language model (LLM) deployments for long-horizon tasks face a fundamental constraint: context windows are finite while productive work sessions are not.
By Shweta Mishra
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations.
arXiv:2607. 21503v1 Announce Type: new Abstract: Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs.
By Gaurav Dadhich
arXiv:2607. 25066v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window.
By Thang Dang, Yuma Ichikawa, Sakina Fatima, Koichi Shirahata
arXiv:2609.37743v1 Announce Type: new
Abstract: LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even...
By Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars
Online web agents frequently add memory, workflow, or skill modules to a base actor, which can boost performance but also consume test‑time tokens—a cost rarely reported. This study evaluates such augmentation under a fixed inference budget, comparing AWM, ASI, and ReasoningBank to a token‑matched vanilla baseline across four WebArena domains and three models (Gemini 3 Flash, GPT‑5.4‑mini, Qwen 3.6‑27B). The vanilla baseline consistently matches or outperforms the augmentation methods in overall success rate while often using fewer tokens, a trend also seen on WorkArena‑L1 with Qwen 3.6‑27B. The results suggest that skills and workflow memory may only be beneficial in specific domains, and that run‑to‑run variance should be reported as a core evaluation criterion for online web agents.
By Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz, Ismail Ben Ayed, Issam H. Laradji, Spandana Gella, Nicolas Gontier