Towards Data Science

Coding Agents Don’t Need Bigger Context Windows — They Need a Context Compiler

Most coding agents treat prompt construction like retrieval: gather more files, add more context, hope the model figures it out. But that approach breaks down fast.

Towards Data Science
Aug 24

AI Agents Don’t Need More Context — They Need Typed Context

The article argues that AI agents face a context typing issue rather than merely a lack of context. It explains how flattening instructions, memory, evidence, and tool outputs into a single string erases semantic boundaries, and presents a lightweight, zero‑dependency Python runtime that preserves these boundaries, tracks provenance, and rejects invalid transformations before they reach the model. The post details the implementation, testing, and the guarantees and limitations of this approach.

By Emmimal P Alexander
arXiv Computation and Language
Sep 23

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.

By Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten
arXiv AI
6d ago

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.

By Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)