VaultGemma: The world's most capable differentially private LLM
We introduce VaultGemma, the most capable model trained from scratch with differential privacy.
We introduce VaultGemma, the most capable model trained from scratch with differential privacy.
arXiv:2607. 06608v1 Announce Type: cross Abstract: We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.
arXiv:2608.28198v1 Announce Type: new Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in...
arXiv:2608.28934v1 Announce Type: new Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In...
arXiv:2606. 04399v1 Announce Type: new Abstract: In the paradigm of decentralized learning, a group of agents collaborate to train a global model using distributed datasets without a central server.
arXiv:2512. 03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users.
arXiv:2512. 04008v2 Announce Type: replace Abstract: Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model.
arXiv:2601. 14660v2 Announce Type: replace-cross Abstract: Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data.
arXiv:2509. 10691v3 Announce Type: replace-cross Abstract: Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sensitive information through inversion, reconstruction, and membership inference attacks.
arXiv:2510. 04902v3 Announce Type: replace Abstract: Tuning hyperparameters in federated machine learning can substantially impact model performance.
PrivateHub is a contrastive diffusion model designed to generate synthetic multi‑sensor data that protects private user activities while keeping non‑private applications detectable. It operates in two stages: App‑Conditioned Pre‑training, which conditions the model on application embeddings, and App‑Aware Fine‑tuning, which uses contrastive learning to separate private from non‑private data. Experiments on three real‑world datasets demonstrate that PrivateHub reduces private‑application inference accuracy by 40–50% without harming non‑private performance and remains robust even when attackers retrain on the synthetic data.
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.