arXiv Machine Learning

Proofs of Ownership for Machine Learning Models

arXiv:2606. 30423v1 Announce Type: new Abstract: With the increasing adoption of Machine Learning, protecting model ownership has become an essential challenge.

arXiv Machine Learning
Sep 10

Characterizing Privacy-Audit Alignment in Behavioral Audit of Machine Unlearning

The paper investigates the privacy risks inherent in auditing machine unlearning (MU) when the audit relies only on querying the model for behavioral signals. It shows that such generic audit schemes inevitably leak information about the retained data set, providing a geometric transfer theorem that bounds the distinguishability of retained set membership based on audit accuracy. The study also analyzes how the unlearned set, target sample, and query protocol influence the privacy‑audit transfer coefficient, with empirical evidence from both convex and non‑convex models supporting the theoretical findings.

By Liou Tang, James Joshi, Ashish Kundu
arXiv Machine Learning
Jul 27

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.

By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
arXiv Machine Learning
Sep 11

SoK: Privacy Attacks on Machine Learning via Explainable AI

The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.

By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
arXiv Machine Learning
Sep 11

CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration

The paper introduces CertDW, a certified dataset watermark and ownership verification method that remains reliable even under malicious perturbations. By leveraging conformal prediction, it defines two statistical measures—principal probability (PP) and watermark robustness (WR)—to evaluate model stability on benign versus watermarked samples. The authors derive certification conditions linking WR to a PP-based threshold and provide a high‑probability bound on false positives, enabling robust ownership verification when a suspicious model’s WR exceeds the PP values of benign models.

By Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao
arXiv AI
Jun 2

Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization

arXiv:2510. 10982v2 Announce Type: replace-cross Abstract: Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications.

By Zihan Wang, Zhiyong Ma, Zhongkui Ma, Shuofeng Liu, Akide Liu, Derui Wang, Minhui Xue, Guangdong Bai
arXiv Machine Learning
Sep 23

Optimizing Canaries for Privacy Auditing with Metagradient Descent

The paper investigates black-box privacy auditing for differentially private learning algorithms, focusing on DP‑SGD. It introduces a method that optimizes the auditor’s canary set using metagradient descent, improving empirical lower bounds on privacy parameters compared to prior canary designs. The approach is shown to be DP‑SGD agnostic and efficient, with optimized canaries for small models remaining effective for larger DP‑SGD models.

By Matteo Boglioni, Terrance Liu, Andrew Ilyas, Zhiwei Steven Wu