SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 02995v1 Announce Type: cross Abstract: Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens.
arXiv:2606. 28479v1 Announce Type: cross Abstract: CSIRTs increasingly fine tune language models on vulnerability scan records, but these records expose internal network topology and create privacy risks under regulations such as GDPR and LGPD.
arXiv:2604.27426v2 Announce Type: replace-cross Abstract: Local fine-tuning datasets routinely contain sensitive secrets such as API keys, personal identifiers, and financial records. Although "local...
arXiv:2606. 14210v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment.
The paper introduces an attack that reconstructs text generated by locally hosted large language models by monitoring CPU cache activity during detokenization. It uses Flush+Reload on shared tokenizer code to time decoding, then Prime+Probe to capture token‑dependent cache traces, followed by a clustering‑and‑language‑model pipeline to recover the output text. The method is evaluated across various datasets, hardware, inference frameworks, and model families, successfully retrieving semantically accurate outputs from real‑world local LLM deployments, including agentic systems.
The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.