An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code...
Paritok-4B is a 4‑billion‑parameter LoRA compressor designed for coding agents, which extracts and retains key spans of code rather than paraphrasing them. It is intent‑conditioned, selecting lines that are most relevant to the agent’s current task, and achieves high fidelity with 96% of identifiers, paths, and numbers preserved. Trained on 67,074 real OpenHands trajectories and fine‑tuned on Qwen3‑4B, it compresses agent context to about 25.7% of its original size while keeping 86.5% of the uncompressed solve quality across 300 SWE‑bench Lite instances.
The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.
KuaFu is a unified behavior‑compression layer that reduces each user behavior item to 2–4 tokens, dramatically shrinking per‑item cache size while preserving fidelity through a four‑stage training process. In production across four profiling tasks, it matches or outperforms uncompressed single‑task models, boosts GPU throughput by 37–350%, and saves 190 GPUs. On public benchmarks it consistently beats prior compressors at the same compression ratio, and on RecBench a 4B KuaFu model outperforms its 8B counterpart by 1.90 points, contributing to a 1.37% lift in overall GMV on Tencent’s advertising and recommendation platform.
arXiv:2607. 15516v1 Announce Type: cross Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent).
arXiv:2609.23790v1 Announce Type: new Abstract: Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected toke...