arXiv AI

Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning

The paper introduces GraphMemory, a lightweight graph-based memory system designed to improve token efficiency in test-time continual learning for large language models. By retrieving only relevant subgraphs for each query, GraphMemory keeps the amount of retrieved memory constant as more examples are processed, avoiding the token cost and performance degradation of traditional shared-context approaches. Experiments demonstrate that GraphMemory achieves competitive downstream performance while using roughly 81‑85% fewer memory‑construction tokens than baseline methods.

arXiv AI
Jul 16

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

arXiv:2607. 13591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks.

By Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu
arXiv AI
Oct 1

Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling

Hermes introduces a family of harnesses that give models control over how they allocate and reuse context windows during inference, a capability termed contextual reasoning. The accompanying Hermes‑Learn framework trains models in two stages to develop these decision‑making skills, enabling them to scale performance with additional compute at test time. Experiments show that while large models naturally benefit, smaller open‑source models can close the performance gap through this training, with gains generalizing across benchmarks, extrapolating beyond trained compute, and transferring to other scaling methods.

By Xinyu Li, Mononito Goswami, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Bl\"obaum, Purak Jain
arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun