arXiv AI By Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Read the original on arXiv AI →

The paper introduces the Correlated Promotion Benchmark (CPB) to evaluate how agents decide whether to admit claims into shared memory, addressing the risk of repeating false claims. CPB offers two modes: CPB-Static, a frozen test set with fixed gold actions, and CPB-Live, which runs multi‑agent teams and tracks source lineage. Experiments across eight admission policies and four agent families show that deduplication reduces false claims but also discards true ones, while gating on declared source type most effectively limits false adoption.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents

The paper introduces the Distributed‑Evidence Paradox, where long‑running LLM agents compress past interactions into persistent memories that may not be fully supported by the interaction history. It defines three key requirements—evidence scope, compositional validity, and admission reliability—and proposes DerivAudit, a framework that checks whether a memory is truly supported by the available history. Experiments on two memory corpora show that expanding the evidence base can recover support for many memories, yet many remain unsupported, and broader evidence alone does not guarantee reliable admission.

By Hongjun Liu, Chen Zhao
arXiv AI
4d ago

Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle

The paper introduces Audience‑Bound Persistent Memory, a system that tracks the audience of each memory item and enforces authorization throughout the memory lifecycle. Each item carries the audience present at recording, and derived items are partitioned or suppressed based on the intersection of source audiences, expanding only through explicit grants. The authors implement the approach in two reference architectures—a flat store and a relationship graph—and evaluate it on 10,000 multi‑party histories, showing that no forbidden items entered any context while unscoped retrieval exposed forbidden items in 82% of cases, and that entitled recall matched policy‑equivalent baselines and outperformed unscoped retrieval by 0.30 Recall@5.

By Sibo Liu
arXiv AI
Sep 3

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.

By Peiying Zhu, Sidi Chang