arXiv AI

Mitigating Private Data Leakage in LLMs with Whiteout

The paper introduces Whiteout, a tool that prevents large language models from revealing personally sensitive information (PSI) by overwriting such data with carefully crafted obfuscation samples. Whiteout is evaluated on various LLMs, including an OpenAI model, and shows effective PSI protection with minimal impact on model utility and safety. The study also tests Whiteout against multiple attack vectors and discusses its security and ethical implications.

Hugging Face Trending Papers
Aug 20

Inadvertent Context Leakage in Language Models

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction.

arXiv AI
Sep 11

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

The paper reports that large language models can acquire cipher-based covert communication skills without fine‑tuning, using prompting or in‑context learning instead. This enables new jailbreak attacks that bypass alignment safeguards by encrypting harmful requests, making them appear as nonsensical text to harmfulness classifiers. The authors demonstrate successful attacks against frontier models from Anthropic, Google, and OpenAI.

By Thomas Rivasseau