arXiv AI By Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang

HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents

Read the original on arXiv AI →

The paper introduces Hard-Origin Adaptively Softened Memory (HasMem), a memory system for large language model agents that combines frozen hard‑prompt embeddings with a controller, writer, reader, and global module to adaptively resize and re‑encode memory entries. On a reconstruction probe of 535 questions, HasMem achieves a lexical F1 of 95.3, outperforming the hard reference by 4.4 percentage points while maintaining 93.6% of the reference’s memory positions. Across six configurations with similar per‑question budgets, the system surpasses rule‑based re‑encoding by 8.0–23.6 exact‑match points, and on LongMemEval‑S it improves local lexical F1 from 3.4 to 8.9 and reduces answer negative log‑likelihood from 12.257 to 5.274.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.

By Yuan Si, Simeng Han, Daming Li, Jialu Zhang
arXiv Computation and Language
Aug 27

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

The paper introduces SCALE-QA, a new QA benchmark that tests conversational memory in flat, unsegmented multi‑topic threads by requiring agents to infer which earlier episode supports a later task decision. The dataset contains 3,000 audited questions across ten domains, uses deterministic four‑way multiple‑choice grading, and includes a runtime builder for reproducibility. The authors also propose Temporal‑Semantic Interleaved Memory Reconstruction (TSIM), a hierarchical memory stack that segments turns into coherent episodes and indexes them with deterministic summaries and cluster‑routing views, achieving significant accuracy gains over strong RAG baselines and long‑context LLMs.

By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie