arXiv:2606. 14518v1 Announce Type: new Abstract: The removal of learned data from Machine Learning models through Machine Unlearning (MU) has been widely studied; however, there has yet to be an agreed-upon scheme for auditing MU.
By Liou Tang, James Joshi, Ashish Kundu
arXiv:2608.28929v1 Announce Type: cross
Abstract: Large-scale diffusion models have fueled numerous profitable downstream applications for AI-related businesses, including visual editing and content...
By Feng Jiang, Zuobin Xiong, An Huang, Zhipeng Cai, Yingshu Li
arXiv:2409. 06130v2 Announce Type: replace-cross Abstract: Modern machine learning models require substantial computational resources and data to train, making them valuable intellectual property.
By Aoting Hu, Yanzhi Chen, Renjie Xie, Xinwei Zhang, Wei Xu
The paper investigates the privacy risks inherent in auditing machine unlearning (MU) when the audit relies only on querying the model for behavioral signals. It shows that such generic audit schemes inevitably leak information about the retained data set, providing a geometric transfer theorem that bounds the distinguishability of retained set membership based on audit accuracy. The study also analyzes how the unlearned set, target sample, and query protocol influence the privacy‑audit transfer coefficient, with empirical evidence from both convex and non‑convex models supporting the theoretical findings.
By Liou Tang, James Joshi, Ashish Kundu
arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.
By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.
By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday