arXiv AI

Can Computation from Earlier Problems Help LLMs Solve New Ones?

The paper investigates whether computation from earlier problems can aid large language models (LLMs) in solving subsequent ones. Preliminary experiments show that retained conversation history can both improve and degrade later-turn accuracy. To address this, the authors propose STAIR, a lightweight module that stores key-value pairs from previous responses and redirects new queries to this bank, improving later-turn accuracy by up to 11.67 percentage points across several benchmarks.

arXiv AI
3d ago

When Context Changes: Understanding Update Failures in LLMs

The paper introduces the concept of stale binding, where large language models (LLMs) answer with outdated values despite having newer information in context. It presents Controlled In-Context Memory (CICM), a benchmark to track and test the use of updated information in conversations and agent logs. Experiments on open‑source models reveal an attention drift mechanism that favors old values, and the authors propose a simple attention‑redirecting intervention that largely corrects these errors without harming correct answers.

By Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei
arXiv AI
4d ago

Mnemon: Raw Records, Fast Judgments, Slow Thoughts

Mnemon is a memory agent that stores conversations as raw, dated records and uses a fast System 1 decision model (Jev) to quickly judge the relevance of records, while a slow System 2 LLM plans searches and composes answers. The agent consolidates records into topic timelines and value histories in the background, enabling efficient retrieval without rewriting conversations into structured formats. Experiments show Mnemon achieving high scores on LoCoMo and LongMemEval‑S with low context length and cost, and Jev outperforming LLMs in evidence separation and speed.

By Guangren Wang
arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv AI
Sep 11

What Should an Agent Forget? Separating What Is Stored from What Is Used

The paper introduces RD-Forget, a training‑free framework that separates what a persistent language agent stores from what it uses at answer time. It keeps a source archive of all observations while a query‑conditioned memory view filters evidence relevant to the current question, using a frozen language‑model curator to group facts into semantic slots and preserve multi‑hop relations. The approach employs rate‑distortion principles to stay within a memory budget and demonstrates improvements across conversational memory, knowledge updating, fact consolidation, long‑context reasoning, and personalization tasks.

By Yuhang Li, Yuchen Li
Hugging Face Trending Papers
Jul 29

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification.