arXiv AI By Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao

Mitigating Private Data Leakage in LLMs with Whiteout

Read the original on arXiv AI →

The paper introduces Whiteout, a tool that prevents large language models from revealing personally sensitive information (PSI) by overwriting such data with carefully crafted obfuscation samples. Whiteout is evaluated on various LLMs, including an OpenAI model, and shows effective PSI protection with minimal impact on model utility and safety. The study also tests Whiteout against multiple attack vectors and discusses its security and ethical implications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 20

Inadvertent Context Leakage in Language Models

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction.