arXiv AI By Mengming Li, Ceyu XU, Qijun Zhang, Jiangnan Yu, Xiangfeng Sun, Haohui Mai, Zhiyao Xie

Memory Compression for High-Fanout Agent Sandboxes

Read the original on arXiv AI →

The paper introduces AgentZip, a memory compression system tailored for high‑fanout AI‑agent sandboxes. It exploits template‑relative and cross‑sandbox redundancy, expands compression to any profitable page, and shifts overhead control to restore‑time prefetching. The approach aligns compression with LLM waiting periods, achieving up to 8.7× memory reduction and reducing slowdown to 1.40× while preserving most savings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 28

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

arXiv:2607. 23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency.

By Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv AI
3d ago

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV is a new KV cache compression framework designed for agentic large language model (LLM) inference. It prioritizes cache entries that contribute to action generation, using action-oriented eviction, confidence-driven budget allocation, and page-aware compression to reduce memory usage while preserving accuracy. In long-trace tasks, ActKV retains 98.53% of FullKV’s accuracy using only 25.98% of its peak memory and boosts token and task throughput by 3.97× and 3.58×, respectively.

By Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou