arXiv Computation and Language By Haoxuan Jia, Yang Liu, Yingguang Yang, Yancheng Chen, Chongyang Zhang, Hao Zheng, Qian Li, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Hao Peng, Junyu Lu, Du Cheng, Philip S. Yu, Bin Chong

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

arXiv AI
2d ago

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

The paper introduces Causal Memory Policy (CMP), a framework that identifies the utility of memories in memory‑augmented language models by intervening on retrieval rather than on storage. CMP reserves fixed context slots for memories sampled with known propensities and estimates utility using self‑normalized inverse propensity weighting, providing unbiased estimates and exact variance. Experiments show that CMP improves discrimination between required and non‑required memories and reveals that identified utility alone is insufficient for retention decisions across unseen queries.

By Arman Behnam, Binghui Wang
Hugging Face Trending Papers
Jul 9

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed.

arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra