arXiv Machine Learning By Jiangze Yan, Yi Shen, Wenjing Zhang, Jieyun Huang, Zhaoxiang Liu, Ning Wang, Kai Wang, Shiguo Lian

HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents

Read the original on arXiv Machine Learning →

arXiv:2606. 16285v1 Announce Type: cross Abstract: Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be rewarded or penalized due to downstream tool failures, noisy observations, or reasoning errors rather than its own contribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 9

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed.

arXiv AI
2d ago

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

The paper introduces Causal Memory Policy (CMP), a framework that identifies the utility of memories in memory‑augmented language models by intervening on retrieval rather than on storage. CMP reserves fixed context slots for memories sampled with known propensities and estimates utility using self‑normalized inverse propensity weighting, providing unbiased estimates and exact variance. Experiments show that CMP improves discrimination between required and non‑required memories and reveals that identified utility alone is insufficient for retention decisions across unseen queries.

By Arman Behnam, Binghui Wang
arXiv AI
Sep 3

CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

CHIME introduces a credit‑aware hierarchical memory evolution framework that separates planning and execution experiences into distinct memory banks. By attributing each task outcome to the plan, execution, both, or neither before memorization, CHIME mitigates bias from noisy final outcomes and improves long‑horizon agent planning. Experiments on four benchmarks demonstrate that CHIME outperforms existing training‑based and self‑evolving memory methods, requires fewer memory items, and transfers effectively across backbone models.

By Yongshi Ye, Tian Lan, Feihu Jiang, Muyang Ye, Bin Zhu, Qianghuai Jia, Longyue Wang, Zhao Xu, Weihua Luo, Xiaodong Shi