ACM: Agentic Context Management for Long Horizon Tasks
arXiv:2607. 23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment.
Most coding agents treat prompt construction like retrieval: gather more files, add more context, hope the model figures it out. But that approach breaks down fast.
arXiv:2607. 23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment.
arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although thi...
The article argues that AI agents face a context typing issue rather than merely a lack of context. It explains how flattening instructions, memory, evidence, and tool outputs into a single string erases semantic boundaries, and presents a lightweight, zero‑dependency Python runtime that preserves these boundaries, tracks provenance, and rejects invalid transformations before they reach the model. The post details the implementation, testing, and the guarantees and limitations of this approach.
Most AI memory systems keep the newest information—not the most important. Here's how I used the Ebbinghaus forgetting curve to build a better memory engine for LLMs.
arXiv:2609.00759v1 Announce Type: new Abstract: Large language models (LLMs) increasingly handle in-context learning (ICL) tasks where a long, novel context defines the rules, knowledge, and output s...
arXiv:2510. 00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise records of actions and observations.
LLMs don’t fail because they forget—they fail because they remember too much. As conversations grow, prompts accumulate redundant and low-value tokens, driving up cost and latency while silently degrading output quality.
Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, but they either learn a compression policy that discards evidence or manage context in a layer the agent never sees.
The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.
arXiv:2608. 06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood.
The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.
arXiv:2607. 20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents.