arXiv Computation and Language

Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents

The paper introduces HiCoMER, a framework that manages hierarchical collaborative memory and performs validity-aware retrieval for large language model agents. HiCoMER distinguishes between team and individual memories, updates their validity, and retrieves only those that remain valid, rather than treating all memories as a flat pool. Experiments on two new collaborative QA datasets show that HiCoMER reduces outdated retrieval, preserves current team consensus, and improves downstream question‑answering quality.

arXiv AI
Sep 10

AMA: Adaptive Memory via Multi-Agent Collaboration

The paper introduces AMA, a framework that uses multiple agents—Constructor, Retriever, Judge, and Refresher—to manage memory for large language model agents. AMA’s hierarchical memory design dynamically adjusts retrieval granularity to match task complexity, while the Judge and Refresher ensure relevance, consistency, and timely updates. Experiments on long-context benchmarks show AMA outperforms existing baselines and cuts token usage by about 80% compared to full-context approaches.

By Weiquan Huang, Zixuan Wang, Hehai Lin, Sudong Wang, Bo Xu, Qian Li, Beier Zhu, Linyi Yang, Chengwei Qin
arXiv Computation and Language
Sep 10

ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations

ROAM is a relation‑guided framework for managing atomic memories in long‑term language‑model agents. It classifies atom pairs as independent, equivalent, subsuming, or conflicting, then organizes observations into Primary and Evidence roles, fusing complementary details into compact views. Only Primary views are retrieved for answering, which reduces redundancy and improves answer accuracy by up to 29.8 percentage points across models and settings.

By Jianjie Zheng, Peng Lai, Sijie Cheng, Jiehui Zhao, Lei Yang, Guanhua Chen
arXiv AI
Jul 21

Accurate and Efficient Long-Term Memory for LLM Agents

arXiv:2607. 16211v1 Announce Type: new Abstract: LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment.

By Zicheng Zhao, Xinyang Guo, Luyao Lv, Menghan Wang, Ming Li, Shuaicheng Li
arXiv AI
Sep 11

What Should an Agent Forget? Separating What Is Stored from What Is Used

The paper introduces RD-Forget, a training‑free framework that separates what a persistent language agent stores from what it uses at answer time. It keeps a source archive of all observations while a query‑conditioned memory view filters evidence relevant to the current question, using a frozen language‑model curator to group facts into semantic slots and preserve multi‑hop relations. The approach employs rate‑distortion principles to stay within a memory budget and demonstrates improvements across conversational memory, knowledge updating, fact consolidation, long‑context reasoning, and personalization tasks.

By Yuhang Li, Yuchen Li
arXiv Machine Learning
Sep 11

Evaluating Memory Structure in LLM Agents

The paper introduces StructMemEval, a benchmark designed to assess how well large language model (LLM) agents can organize their long‑term memory rather than merely recall facts. It compiles tasks that humans typically solve by structuring knowledge—such as transaction ledgers, to‑do lists, and trees—and evaluates agents on these. Experiments show that simple retrieval‑augmented LLMs struggle with such organization tasks, while memory‑augmented agents perform better when explicitly prompted to structure their memory, yet many modern LLMs still fail to recognize memory structures without prompting.

By Alina Shutova, Alexandra Olenina, Ivan Vinogradov, Anton Sinitsin