DeepMind Blog

VaultGemma: The world's most capable differentially private LLM

We introduce VaultGemma, the most capable model trained from scratch with differential privacy.

arXiv Machine Learning
Jul 23

Differentially Private Neural Network Training Under the Hidden State Assumption

arXiv:2407. 08233v3 Announce Type: replace Abstract: Current differentially private learning paradigms face a severe utility bottleneck: DP-SGD degrades performance through noise accumulation over training steps, while aggregation-based approaches such as PATE suffer from data inefficiency due to disjoint data partitioning.

By Ding Chen, Haochen Luo, Xiaofei Wang, Chen Liu
arXiv Machine Learning
4d ago

Certification-Based Differentially Private Learning

The paper extends the abstract gradient training (AGT) framework to provide tighter differential privacy guarantees for both private prediction and private learning. It introduces Abstract Gradient Sampling (AGS) to analyze privacy in continuous, unbounded regression and offers theoretical and empirical evidence that these methods yield tighter bounds than global-sensitivity baselines, even in previously unbounded settings. The authors also demonstrate that their private learning algorithm can outperform standard private learners under comparable conditions.

By Mihnea Ghitu, Matthew Wicker
arXiv AI
Sep 4

Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis

The paper introduces a federated inference framework that enables multiple commercial large language model (LLM) APIs—such as LLaMA‑3.3‑70B, GPT‑4o‑mini, and Claude‑3‑Haiku—to collaborate on cognitive diagnosis tasks without accessing raw student data or proprietary model internals. Each entity’s predictions are perturbed with Laplace noise to provide epsilon‑local differential privacy, and a residual‑based aggregation scheme mitigates model heterogeneity. Experiments on three educational benchmarks demonstrate strong privacy guarantees with minimal accuracy loss, confirming the framework’s practical usability and cross‑domain generalizability.

By Yagna Manasa Boyapati, Chong Yu, Tianyu Jiang, Justin Zhan
arXiv AI
Aug 18

Privacy-Preserving Decentralized Federated Learning via Explainable Adaptive Differential Privacy

arXiv:2509. 10691v3 Announce Type: replace-cross Abstract: Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sensitive information through inversion, reconstruction, and membership inference attacks.

By Fardin Jalil Piran, Zhiling Chen, Yang Zhang, Qianyu Zhou, Jiong Tang, Farhad Imani
arXiv Machine Learning
Jun 30

Efficient Unlearning with Privacy Guarantees

arXiv:2507. 04771v2 Announce Type: replace-cross Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them.

By Josep Domingo-Ferrer, Najeeb Jebreel, David S\'anchez
arXiv AI
Sep 15

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.

By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar