arXiv AI

Privacy-Preserving AI Verification via Minimal Information Disclosure

arXiv:2608. 02774v1 Announce Type: cross Abstract: AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workload, or hardware.

arXiv AI
2d ago

Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

Tokenized Key-Gated Adapter Routing (Locket) is a framework that embeds fine‑grained, policy‑driven access control into large language models by training lightweight LoRA adapters for different privacy policies. A gating module associates a learned keyed entry token with a specific adapter, allowing authorized tokens to unlock private knowledge while invalid or missing tokens trigger privacy‑preserving adapters that redact or sanitize sensitive content. Experiments on datasets such as Enron, ECHR, and Yelp with models like Qwen3, Llama‑3.2, and Gemma‑2‑2B show that Locket maintains perplexity comparable to fine‑tuning when the correct token is provided, and significantly reduces PII leakage when the token is absent or invalid, without sacrificing utility.

By Mohamed Shaaban, Mohamed Elmahallawy
arXiv Machine Learning
Sep 11

SoK: Privacy Attacks on Machine Learning via Explainable AI

The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.

By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
arXiv AI
Jun 6

Zero knowledge verification for frontier AI training is possible

arXiv:2606. 05433v1 Announce Type: new Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists.

By Pierre Peign\'e, Ky Nguyen, Paul Wang