arXiv Machine Learning

Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance

arXiv AI
Jun 16

Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations

arXiv:2606. 15903v1 Announce Type: cross Abstract: Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the control plane that mutates them via supersede, release, purge (largely untested) -- shapes which forgetting failure modes the system recovers.

By Dongxu Yang
arXiv Computation and Language
Aug 25

Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whet...

By Ayush Gupta, Hima Varshini Surisetty, Sreevidya Bollineni, Varad Ingale, Tuhina Tripathi, Abhishek Lalwani, Somya Chatterjee, Sadid Hasan
arXiv Machine Learning
Sep 16

Test-Time Unlearning via Sparse Autoencoder

arXiv:2609.16229v1 Announce Type: new Abstract: Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify...

By Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Sijia Liu, Tianlong Chen
arXiv AI
Jun 2

SORA: Free Second-Order Attacks in Fast Adversarial Training

arXiv:2606. 00738v1 Announce Type: cross Abstract: Adversarial Training (AT) is a leading defense against adversarial examples but often suffers from Catastrophic Overfitting (CO) in efficient single-step variants, where robustness to multi-step attacks collapses despite high single-step performance.

By Mazdak Teymourian, Ramtin Moslemi, Farzan Rahmani, Mohammad Hossein Rohban
arXiv Machine Learning
Sep 3

Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning

The paper investigates whether a model that has undergone class unlearning can still recover forgotten classes without access to original data. It introduces a white‑box audit method that generates synthetic probes in representation space, filters them by confidence, and relabels boundary‑adjacent probes as the forgotten class. The authors define a Relearning Score to quantify recovery while preserving retain performance, and demonstrate that several unlearning techniques on CIFAR‑10, CIFAR‑100, and TinyImageNet can be fully recovered in a source‑free setting, sometimes even outperforming a retrained reference.

By Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri