arXiv AI

ACM: Agentic Context Management for Long Horizon Tasks

arXiv:2607. 23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment.

arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
arXiv AI
Jul 1

ACE: Pluggable Adaptive Context Elasticizer across Agents

arXiv:2606. 31564v1 Announce Type: new Abstract: The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows.

By Ning Liao, Zihao Long, Xiaoxing Wang, Xue Yang, Yaoming Wang, Ziyuan Zhuang, Xunliang Cai, Rongxiang Weng, Junchi Yan
arXiv Computation and Language
Sep 23

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.

By Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten
arXiv AI
4d ago

Context Language Models

The paper introduces Context Language Models (CLMs), which treat context as a mutable file that the model can update freely, enabling the model to learn what information to retain. CLMs built zero‑shot from existing models outperform state‑of‑the‑art context‑management methods on several benchmarks, achieving higher accuracy with fewer FLOPs. The authors also demonstrate that CLMs can be steered via natural‑language instructions and online reinforcement learning, and they propose a suffix‑cache reuse strategy that further reduces server‑side compute.

By Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh
arXiv AI
4d ago

FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

The paper introduces FOCUS, a training‑free framework that compresses the interaction history of large language model agents by preserving only the past interactions that causally influence future decisions. Unlike prior methods that learn compression policies offline, FOCUS operates entirely at test time, requiring no additional data collection or fine‑tuning and can be applied to any closed‑API model. Experiments on a variety of agentic benchmarks show that FOCUS reduces peak context length by up to 48% and dependency by 73%, while improving task success by up to 8.9 percentage points.

By Shantanu Dixit, Anson Bastos, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan