arXiv AI By Thomas Rivasseau

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

Read the original on arXiv AI →

The paper reports that large language models can acquire cipher-based covert communication skills without fine‑tuning, using prompting or in‑context learning instead. This enables new jailbreak attacks that bypass alignment safeguards by encrypting harmful requests, making them appear as nonsensical text to harmfulness classifiers. The authors demonstrate successful attacks against frontier models from Anthropic, Google, and OpenAI.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

HE-Guardrail is a framework that applies homomorphic encryption to enforce guardrails against jailbreak attacks during encrypted large language model inference. It evaluates guardrail mechanisms—Llama Guard, JBShield, and GradSafe—directly on encrypted data, deciding whether to return the model’s response to the client. The approach preserves confidentiality while closely matching the decisions of plaintext guardrails, offering varied security-efficiency-utility trade‑offs.

By Byeongseo Min, Yongwoo Lee, Young-Sik Kim, Yongjune Kim