arXiv Computation and Language
Aug 27

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

The paper introduces the Groundhog Bit-Flip Attack (GBFA), a novel denial-of-service attack targeting Mixture-of-Experts (MoE) large language models (LLMs). By flipping specific routing-layer bits that activate certain experts, GBFA can cause models to generate excessively long outputs—up to a 5912% increase in token usage—while largely preserving semantic content. The attack requires deactivating fewer than four experts on average across four real-world MoE-based LLMs, exposing a significant robustness vulnerability in these architectures.

By Huakang Lin, Tiancheng Zheng, Mingxuan Sun, Tianhong Xu, Fan Zhang, Yunsi Fei, Ruyi Ding