arXiv AI By Yuan Si, Simeng Han, Daming Li, Jialu Zhang

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

Read the original on arXiv AI →

RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents

The paper introduces Hard-Origin Adaptively Softened Memory (HasMem), a memory system for large language model agents that combines frozen hard‑prompt embeddings with a controller, writer, reader, and global module to adaptively resize and re‑encode memory entries. On a reconstruction probe of 535 questions, HasMem achieves a lexical F1 of 95.3, outperforming the hard reference by 4.4 percentage points while maintaining 93.6% of the reference’s memory positions. Across six configurations with similar per‑question budgets, the system surpasses rule‑based re‑encoding by 8.0–23.6 exact‑match points, and on LongMemEval‑S it improves local lexical F1 from 3.4 to 8.9 and reduces answer negative log‑likelihood from 12.257 to 5.274.

By Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang
arXiv Computation and Language
Aug 27

AWM: Answerable Working Memory for Long-Document VQA Agents

The paper introduces AWM, a framework that treats the terminal working memory of long‑document VQA agents as an answerable evidence artifact. It proposes a memory‑only answerability diagnostic and incorporates this signal into the GRPO reward, giving higher advantage to trajectories whose final memory can answer the question alone. Experiments on MMLongBench‑Doc and LongDocURL show that AWM‑GRPO boosts final‑answer accuracy by up to 11.9 points and reduces the rate of correct answers that cannot be supported by memory alone.

By Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang, Rui Lu, Yuxiao Dong, Jie Tang, Evgeny Kharlamov