arXiv AI By Olukunle Owolabi, Pulkit Gupta, Fei Wang

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

Read the original on arXiv AI →

The paper proposes a category‑conditioned retention strategy for agent memory, arguing that a single global confidence threshold is insufficient to balance reliability across different assertion types. By applying stricter retention bars to value assertions while allowing more liberal retention for other categories, the authors demonstrate a 36% relative reduction in unsupported value claims and a 13‑percentage‑point increase in coverage compared to a global threshold. The study is evaluated on a cold‑start memory pipeline with 100 synthetic personas, showing that selective, category‑aware retention improves overall memory reliability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

The paper presents a retrieval‑admissibility verification framework for long‑term memory agents, classifying each memory‑query pair as admissible, inadmissible, or unresolved. It evaluates the framework on public benchmarks (RHELM and MemOps), showing improved anchor recall and reduced exact similarity errors, while also revealing that existing verifiers miss certain inadmissible exposures. The study highlights the need for separate checks on candidate support, admissibility, prompt exposure, and answer disclosure to ensure safe memory retrieval.

By Zi Wang, Xingqiao Wang, Emmanuel Addai, Devika Ambekar, Xiaowei Xu
arXiv AI
Jun 16

Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations

arXiv:2606. 15903v1 Announce Type: cross Abstract: Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the control plane that mutates them via supersede, release, purge (largely untested) -- shapes which forgetting failure modes the system recovers.

By Dongxu Yang
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.