arXiv AI By Haorui Chen, Yuancheng Zhu, Yitong Zhang, Jia Li

CoACT: Action-Preserving Observation Compression for Coding Agents

Read the original on arXiv AI →

arXiv:2607. 02911v1 Announce Type: cross Abstract: LLM-based coding agents solve software-engineering tasks through iterative interactions with development environments, where returned observations accumulate in the context and become a major source of inference cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

The paper introduces Dynamic Tool Output Compression (DTOC), a framework that manages context in large‑language‑model agents by storing full tool outputs in external memory and inserting compact placeholders into the active context. DTOC treats context updates as explicit, reversible operations within the agent’s reasoning loop, allowing selective reconstruction of compressed outputs when needed. Experiments on the DeepSWE benchmark show that for responsive models such as Sonnet 4.6 and GPT‑5.4, DTOC reduces input tokens and agent steps while significantly improving solve rates and lowering cost per solved task, with ablation studies confirming the importance of reversibility for maintaining performance.

By Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten
arXiv AI
Sep 25

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Long‑horizon language model agents accumulate reasoning history, which inflates context length and inference cost. The paper introduces Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training‑free online method that ranks and removes reasoning blocks based on frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR raises average reward from 0.699 to 0.718 and cuts input, output, and cache read tokens by 25.5%, 14.4%, and 33.3% respectively, while analyses show that historical reasoning becomes replaceable once task‑relevant state is externalized.

By Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv AI
Sep 28

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV is a new KV cache compression framework designed for agentic large language model (LLM) inference. It prioritizes cache entries that contribute to action generation, using action-oriented eviction, confidence-driven budget allocation, and page-aware compression to reduce memory usage while preserving accuracy. In long-trace tasks, ActKV retains 98.53% of FullKV’s accuracy using only 25.98% of its peak memory and boosts token and task throughput by 3.97× and 3.58×, respectively.

By Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou