Adversarial Entropy Inflation Against Gumbel-Based Inference Verification
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 05958v1 Announce Type: new Abstract: Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning.
arXiv:2606. 09135v1 Announce Type: cross Abstract: We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model weights, sampling code, or output distributions.
arXiv:2512. 21815v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs) achieve remarkable performance but remain vulnerable to adversarial attacks.
The paper introduces the Groundhog Bit-Flip Attack (GBFA), a novel denial-of-service attack targeting Mixture-of-Experts (MoE) large language models (LLMs). By flipping specific routing-layer bits that activate certain experts, GBFA can cause models to generate excessively long outputs—up to a 5912% increase in token usage—while largely preserving semantic content. The attack requires deactivating fewer than four experts on average across four real-world MoE-based LLMs, exposing a significant robustness vulnerability in these architectures.
arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.
arXiv:2510. 01529v3 Announce Type: replace Abstract: Ball et al.