Towards Encrypted Large Language Models with FHE
Related stories
Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
arXiv:2310. 16152v5 Announce Type: replace-cross Abstract: Federated learning (FL) has become a key component in various language modeling applications such as machine translation, next-word prediction, and medical record analysis.
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
The paper reports that large language models can acquire cipher-based covert communication skills without fine‑tuning, using prompting or in‑context learning instead. This enables new jailbreak attacks that bypass alignment safeguards by encrypting harmful requests, making them appear as nonsensical text to harmfulness classifiers. The authors demonstrate successful attacks against frontier models from Anthropic, Google, and OpenAI.
HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation
HEAT introduces a fine‑tuning method that treats the number of iterations used to approximate nonlinearities in fully homomorphic encryption (FHE) as learnable parameters, allowing them to co‑adapt with model weights. By optimizing iteration counts per nonlinearity, HEAT reduces the required iterations, bootstraps, and overall latency for encrypted GPT‑2 decoding while improving decode agreement. The approach achieves a 3.1× reduction in iterations, a 1.6× reduction in bootstraps, and a 1.4× speed‑up in end‑to‑end latency without changing the model architecture or requiring retraining from scratch.
On the Recoverability of Private Information Unlearning in Large Language Models
arXiv:2608.29943v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to re...
SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models
arXiv:2606. 05004v1 Announce Type: cross Abstract: With the widespread deployment of public large language models (LLMs) such as ChatGPT, protecting user prompt privacy has become an increasingly critical issue.
Natural Identifiers for Privacy and Data Audits in Large Language Models
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
HE-Guardrail is a framework that applies homomorphic encryption to enforce guardrails against jailbreak attacks during encrypted large language model inference. It evaluates guardrail mechanisms—Llama Guard, JBShield, and GradSafe—directly on encrypted data, deciding whether to return the model’s response to the client. The approach preserves confidentiality while closely matching the decisions of plaintext guardrails, offering varied security-efficiency-utility trade‑offs.
Bypassing Prompt Guards in Production with Controlled-Release Prompting
arXiv:2510. 01529v3 Announce Type: replace Abstract: Ball et al.
Detecting and Understanding Vulnerabilities in Fully Homomorphic Encryption Frameworks
Fully homomorphic encryption (FHE) allows computations to be performed directly on encrypted data without decryption, offering strong privacy guarantees for sensitive data analysis. This capability is important for privacy-sensitive applications like secure cloud computing, finance, and healthcare.
Privacy from Symmetry: Orthogonally Equivariant Transformers for LLM Inference
arXiv:2606. 16461v1 Announce Type: new Abstract: Running large language models locally is often impractical, pushing inference on sensitive text to third-party providers.
VaultGemma: The world's most capable differentially private LLM
We introduce VaultGemma, the most capable model trained from scratch with differential privacy.