arXiv Machine Learning By Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Read the original on arXiv Machine Learning →

CliffCompaction is an autocompaction technique that reduces cost by up to 50% while maintaining or improving performance on benchmarks such as Terminal‑Bench and KernelBench. It achieves this by truncating or dropping content without rephrasing, ensuring compacted information remains faithful and preventing context drift. The method enables efficient test‑time scaling, matching or surpassing higher‑cost models like Opus 4.7 and GPT‑5.3 Codex, and delivers significant CUDA kernel speedups on KernelBench.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 10

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents

arXiv:2607. 08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and storing what happened in previous turns and sessions.

By Ashwin Gerard Colaco, Nada Lahjouji
arXiv AI
3d ago

SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation

SCLATE is a new execution substrate that allows continual‑learning benchmarks and agents to share a single event scheduler via adapters, enabling tasks, session events, and memory consolidation to run on a compressed, real‑time timeline. It also functions as a rollout engine that records every model call’s tokens and log probabilities without modifying the agent’s harness or memory. Using SCLATE, the authors ported seven benchmarks, compared ten harness‑memory configurations across ten models, and demonstrated that post‑training Qwen3.5‑4B can effectively leverage both harness and memory, improving performance on multiple metrics.

By Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu
arXiv Machine Learning
Sep 7

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

KVMem is a KV-context virtualization system that allows large language model agents to maintain workspaces exceeding both GPU key‑value capacity and the model’s native context window. It stores overflowed history as paged KV state across GPU memory, host memory, and NVMe, using lightweight, model‑native attention‑space indexes to retrieve relevant historical blocks. Evaluations on long‑context agent benchmarks show that KVMem improves task utility and inference efficiency, enabling up to one million‑token workspaces on consumer GPUs and achieving interactive responsiveness in local deployments.

By Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu
Hugging Face Trending Papers
Aug 6

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration.