arXiv AI By Christopher J. Chanhnourack

Auditing Long-Term Memory Evaluation: Repeated Judging, Reader Variation, and Negative Controls

Read the original on arXiv AI →

The report audits a long‑term‑memory retrieval chain on 500 LongMemEval‑S development questions, evaluating reader performance under an adapted GPT‑4o rubric. Scores vary widely across reader lanes (93–479), with the strongest lanes scoring 479 and 475, and re‑judging the same pass‑1 answers changes three labels, yielding 478. Fixed‑answer knowledge‑update re‑scoring achieves 70/72 or 69/72 depending on the template, and a different‑family reader scores 474, just 1.0 percentage point below the headline pass. The study also reports that live reader request bodies were not retained, that 18 of 23 gains occurred where baseline packets lacked evidence, and that a negative control rejects a verifier that repairs some wrong drafts but breaks many correct ones. Overall, the findings do not establish a new leaderboard leader or a transferable memory advantage, and the released artifacts support packet inspection and re‑scoring but do not reconstruct the method.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper reports a preregistered audit of language‑model judges used as measurement instruments, revealing that the assumption that a model’s responses remain stable over time is invalid. Across nearly 53,000 audited requests, repeat rankings and byte‑identical replays fell far below required reliability thresholds, with three identified mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—explaining the discrepancy. The study proposes a three‑level snapshot‑identity framework, eight design rules, and a reporting checklist to prevent such reliability failures in future evaluations.

By Haoyaun Zhu, Jie Zhang
arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv AI
Sep 16

LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

The paper introduces LSREP, a Longitudinal State‑Replay Evaluation Protocol designed to assess how conversational memory evolves over time, incorporating ordered replay, lifecycle schedules, repeated probes, evolving reference answers, and mechanism‑fidelity checks. It applies LSREP to ICE v2, a local‑first memory middleware, and reports that on three ordinary‑density datasets ICE v2 achieves near‑zero mean quality difference from vector‑RAG while using fewer fragments but slightly more prompt tokens, yet fails catastrophically on a dense dataset. In a public diagnostic, ICE v2 underperforms pure vector‑RAG on LongMemEval, revealing significant multi‑session and temporal failures and a quality‑cost trade‑off rather than superior efficiency.

By Deepesh Sonar
arXiv AI
2d ago

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

The paper presents a retrieval‑admissibility verification framework for long‑term memory agents, classifying each memory‑query pair as admissible, inadmissible, or unresolved. It evaluates the framework on public benchmarks (RHELM and MemOps), showing improved anchor recall and reduced exact similarity errors, while also revealing that existing verifiers miss certain inadmissible exposures. The study highlights the need for separate checks on candidate support, admissibility, prompt exposure, and answer disclosure to ensure safe memory retrieval.

By Zi Wang, Xingqiao Wang, Emmanuel Addai, Devika Ambekar, Xiaowei Xu