The paper investigates how to provide verifiable deletion certificates for a dense key–value context memory used in support‑vector‑based readouts. By assigning explicit weights to keys and using a one‑class support‑vector boundary, the authors show that reserve keys can be removed without re‑solving, while active keys can be deleted with a decremental solver that matches the result of a full re‑solve. Extensive experiments on synthetic, near‑duplicate, clinical, and learned key sets demonstrate that maintained deletion achieves the same reference state as re‑solve, with negligible readout disagreement and significant speedups.
By Vishwajith Ramesh
The paper investigates whether a model that has undergone class unlearning can still recover forgotten classes without access to original data. It introduces a white‑box audit method that generates synthetic probes in representation space, filters them by confidence, and relabels boundary‑adjacent probes as the forgotten class. The authors define a Relearning Score to quantify recovery while preserving retain performance, and demonstrate that several unlearning techniques on CIFAR‑10, CIFAR‑100, and TinyImageNet can be fully recovered in a source‑free setting, sometimes even outperforming a retrained reference.
By Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it emits that stale value as a confident answer; give the same model an empty memory and it abstains.
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
By Qing Ye, Meng-Hsuan Lin
arXiv:2609.26242v1 Announce Type: new
Abstract: When a data license expires, deleting stored records does not remove influence encoded in a trained forecaster. Machine unlearning seeks to remove this...
By Junyi Ye
arXiv:2608. 11822v1 Announce Type: cross Abstract: A growing body of work reports that language models represent task-relevant latent structure that they fail to use.
By Xining Xun