arXiv AI

When Context Changes: Understanding Update Failures in LLMs

The paper introduces the concept of stale binding, where large language models (LLMs) answer with outdated values despite having newer information in context. It presents Controlled In-Context Memory (CICM), a benchmark to track and test the use of updated information in conversations and agent logs. Experiments on open‑source models reveal an attention drift mechanism that favors old values, and the authors propose a simple attention‑redirecting intervention that largely corrects these errors without harming correct answers.

arXiv AI
Sep 2

Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning

The paper introduces the Unified Memory Agent (UMA), a system that builds a query‑agnostic external memory from a data stream and reuses it across multiple question‑answering sessions. UMA employs a single policy to manage a structured Memory Bank via CRUD operations and uses Task‑Stratified GRPO to supervise memory maintenance based on QA trajectory rewards. The authors also present Ledger‑QA, a benchmark for long‑horizon state tracking, and demonstrate that UMA outperforms other methods on test‑time learning and accurate‑retrieval tasks, with UMA‑Specialist further improving performance after task adaptation.

By Kehao Zhang, Shangtong Gui, Sheng Yang, Wei Chen, Yang Feng
arXiv AI
3d ago

Can Computation from Earlier Problems Help LLMs Solve New Ones?

The paper investigates whether computation from earlier problems can aid large language models (LLMs) in solving subsequent ones. Preliminary experiments show that retained conversation history can both improve and degrade later-turn accuracy. To address this, the authors propose STAIR, a lightweight module that stores key-value pairs from previous responses and redirects new queries to this bank, improving later-turn accuracy by up to 11.67 percentage points across several benchmarks.

By Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
arXiv AI
Aug 24

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

The paper introduces the Weighted Memory Tree (WMT), a hierarchical memory system for large language model agents that organizes execution histories into tasks, subtasks, and actions while assigning each memory a dynamic retention score. Event-based updates and selection-based decay allow WMT to preserve useful information, fold completed trajectories, suppress low-utility content, and retain access to folded context. Experiments on GAIA-Text with Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B show that WMT improves accuracy by an average of 9.97 percentage points and reduces prompt-token usage by 32.8%, while also limiting the persistence of unreliable information.

By Quang Dao, Purvi Kathalkar, Kenneth Eaton