Paritok-4B is a 4‑billion‑parameter LoRA compressor designed for coding agents, which extracts and retains key spans of code rather than paraphrasing them. It is intent‑conditioned, selecting lines that are most relevant to the agent’s current task, and achieves high fidelity with 96% of identifiers, paths, and numbers preserved. Trained on 67,074 real OpenHands trajectories and fine‑tuned on Qwen3‑4B, it compresses agent context to about 25.7% of its original size while keeping 86.5% of the uncompressed solve quality across 300 SWE‑bench Lite instances.
By Jiayu Shi, Luzhuo Chen
arXiv:2608.21690v1 Announce Type: new
Abstract: LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier i...
By Yin Lin, Elaine Ang, Erkang Zhu, Bolin Ding, Jingren Zhou
The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.
By Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten
Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, but they either learn a compression policy that discards evidence or manage context in a layer the agent never sees.
LatentPress compresses conversational histories and long documents into continuous memory tokens that a frozen decoder can read directly, eliminating the need for text reconstruction at inference. The method achieves 4–16× compression with only a small adapter (0.1% of the decoder’s parameters) and outperforms text summaries and OCR-based compression on LongMemEval and LongBench-QA benchmarks. Writing and reading are significantly faster than traditional text summarization or OCR reconstruction, demonstrating a practical machine-facing context interface beyond text and vision.
By Zhengze Zhou, Hejian Sang
Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code...