arXiv AI

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.

Hugging Face Trending Papers
Aug 20

Inadvertent Context Leakage in Language Models

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction.

arXiv AI
Sep 25

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

PrivDrift is a benchmark that tests whether user‑disclosed secrets can still be recovered by large language models after the conversation shifts to unrelated topics. It includes 1,000 controlled multi‑turn dialogues with seeded secrets, topic‑drift turns, and standardized extraction probes. Experiments on three LLMs with extended context windows show that dialogue‑level leakage remains substantial—between 38.7% and 54.6%—and is influenced by model, secret type, and persuasion intensity, while additional topic drift does not reliably reduce leakage.

By Luciano Maldonado
arXiv Computation and Language
Sep 11

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.

By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
Hugging Face Trending Papers
Sep 10

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard identifies when a target model’s behavior shifts toward attacker‑controlled outputs, a signal that appears whenever a backdoor is triggered. The method works across various backdoor types and model families, reliably detecting stealthy attacks that bypass input‑level filters while avoiding the extra generation cost of existing runtime detectors.