arXiv AI
3d ago

When Context Changes: Understanding Update Failures in LLMs

The paper introduces the concept of stale binding, where large language models (LLMs) answer with outdated values despite having newer information in context. It presents Controlled In-Context Memory (CICM), a benchmark to track and test the use of updated information in conversations and agent logs. Experiments on open‑source models reveal an attention drift mechanism that favors old values, and the authors propose a simple attention‑redirecting intervention that largely corrects these errors without harming correct answers.

By Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei
arXiv AI
Sep 2

Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning

The paper introduces the Unified Memory Agent (UMA), a system that builds a query‑agnostic external memory from a data stream and reuses it across multiple question‑answering sessions. UMA employs a single policy to manage a structured Memory Bank via CRUD operations and uses Task‑Stratified GRPO to supervise memory maintenance based on QA trajectory rewards. The authors also present Ledger‑QA, a benchmark for long‑horizon state tracking, and demonstrate that UMA outperforms other methods on test‑time learning and accurate‑retrieval tasks, with UMA‑Specialist further improving performance after task adaptation.

By Kehao Zhang, Shangtong Gui, Sheng Yang, Wei Chen, Yang Feng