arXiv AI

In Defense of Information Leakage in Concept-based Models

arXiv:2606. 10669v1 Announce Type: cross Abstract: Concept-based models (CMs), deep neural networks that ground their predictions on representations aligned with human-understandable concepts (e.

Hugging Face Trending Papers
Aug 20

Inadvertent Context Leakage in Language Models

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction.