Hugging Face Blog

Towards Encrypted Large Language Models with FHE

arXiv AI
Sep 11

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

The paper reports that large language models can acquire cipher-based covert communication skills without fine‑tuning, using prompting or in‑context learning instead. This enables new jailbreak attacks that bypass alignment safeguards by encrypting harmful requests, making them appear as nonsensical text to harmfulness classifiers. The authors demonstrate successful attacks against frontier models from Anthropic, Google, and OpenAI.

By Thomas Rivasseau
arXiv AI
Sep 3

HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation

HEAT introduces a fine‑tuning method that treats the number of iterations used to approximate nonlinearities in fully homomorphic encryption (FHE) as learnable parameters, allowing them to co‑adapt with model weights. By optimizing iteration counts per nonlinearity, HEAT reduces the required iterations, bootstraps, and overall latency for encrypted GPT‑2 decoding while improving decode agreement. The approach achieves a 3.1× reduction in iterations, a 1.6× reduction in bootstraps, and a 1.4× speed‑up in end‑to‑end latency without changing the model architecture or requiring retraining from scratch.

By Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos, Giuseppe Ateniese, Emanuele Rodol\`a
arXiv AI
Sep 21

HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

HE-Guardrail is a framework that applies homomorphic encryption to enforce guardrails against jailbreak attacks during encrypted large language model inference. It evaluates guardrail mechanisms—Llama Guard, JBShield, and GradSafe—directly on encrypted data, deciding whether to return the model’s response to the client. The approach preserves confidentiality while closely matching the decisions of plaintext guardrails, offering varied security-efficiency-utility trade‑offs.

By Byeongseo Min, Yongwoo Lee, Young-Sik Kim, Yongjune Kim