arXiv AI

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

arXiv:2607. 19442v1 Announce Type: cross Abstract: Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes.

arXiv Machine Learning
Sep 10

What a Deletion Certificate Covers, and Where It Expires: Auditable Removal from a Support-Vector Memory

The paper investigates how to provide verifiable deletion certificates for a dense key–value context memory used in support‑vector‑based readouts. By assigning explicit weights to keys and using a one‑class support‑vector boundary, the authors show that reserve keys can be removed without re‑solving, while active keys can be deleted with a decremental solver that matches the result of a full re‑solve. Extensive experiments on synthetic, near‑duplicate, clinical, and learned key sets demonstrate that maintained deletion achieves the same reference state as re‑solve, with negligible readout disagreement and significant speedups.

By Vishwajith Ramesh
arXiv Machine Learning
Sep 3

Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning

The paper investigates whether a model that has undergone class unlearning can still recover forgotten classes without access to original data. It introduces a white‑box audit method that generates synthetic probes in representation space, filters them by confidence, and relabels boundary‑adjacent probes as the forgotten class. The authors define a Relearning Score to quantify recovery while preserving retain performance, and demonstrate that several unlearning techniques on CIFAR‑10, CIFAR‑100, and TinyImageNet can be fully recovered in a source‑free setting, sometimes even outperforming a retrained reference.

By Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin