arXiv Machine Learning By Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

Read the original on arXiv Machine Learning →

arXiv:2607. 20723v1 Announce Type: cross Abstract: This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

The paper introduces an attack that reconstructs text generated by locally hosted large language models by monitoring CPU cache activity during detokenization. It uses Flush+Reload on shared tokenizer code to time decoding, then Prime+Probe to capture token‑dependent cache traces, followed by a clustering‑and‑language‑model pipeline to recover the output text. The method is evaluated across various datasets, hardware, inference frameworks, and model families, successfully retrieving semantically accurate outputs from real‑world local LLM deployments, including agentic systems.

By Roy Weiss, Benyamin Konstantinov, Eitam Sheetrit, Tomer Simon, Yisroel Mirsky
arXiv Computation and Language
Sep 11

SpecGuard: Inference-Time Backdoor Detection For Free

SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.

By Rui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich, Zheng Li
Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.