arXiv AI

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

The paper introduces Direct Relational Set‑Risk Pruning (DRSR), a method for compressing the history of long‑horizon language‑model agents by selecting deletion sets based on risk constraints rather than independent unit scores. DRSR builds counterfactual supervision offline, then uses a lightweight scorer to predict set‑level harm during deployment, removing the largest safe set while respecting recency, protocol, and budget limits. Experiments on WorkBuddyBench Full260 and Eval40 show that DRSR improves mean reward from 0.699 to 0.802 and reduces token usage by over 20%, with further analyses highlighting the importance of decision‑conditioned relations, retained context, pair interactions, and abstention.

arXiv Machine Learning
Sep 10

Can an AI Assistant Really Forget? Auditable Deletion from Addressable Memory

This paper introduces a deletion interface for a pretrained language model, measuring how effectively deleted records are removed from the model’s memory. By retrofitting a support‑vector memory gate into the global attention layers of a frozen Gemma 3, the authors show that deletions can be performed without altering weights and that the resulting state is close to a reference state that never stored the record. Experiments on 4B‑parameter models demonstrate low perplexity impact and strong evidence that deleted content is hard to recover, while larger or smaller models fail to achieve the same guarantees.

By Vishwajith Ramesh
arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv AI
Jul 24

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

arXiv:2607. 21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training.

By Chaofan Pan, Lingfei Ren, Xiangyu Jiang, Yanhua Li, Xuemei Cao, Xiangkun Wang, Hao Yu, Wei Wei, Xin Yang
arXiv AI
Sep 18

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.

By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie