arXiv Computation and Language By Huakang Lin, Tiancheng Zheng, Mingxuan Sun, Tianhong Xu, Fan Zhang, Yunsi Fei, Ruyi Ding

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

Read the original on arXiv Computation and Language →

The paper introduces the Groundhog Bit-Flip Attack (GBFA), a novel denial-of-service attack targeting Mixture-of-Experts (MoE) large language models (LLMs). By flipping specific routing-layer bits that activate certain experts, GBFA can cause models to generate excessively long outputs—up to a 5912% increase in token usage—while largely preserving semantic content. The attack requires deactivating fewer than four experts on average across four real-world MoE-based LLMs, exposing a significant robustness vulnerability in these architectures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
1d ago

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

The paper introduces SEAL, a training-time, parameter‑efficient defense that attaches a plug‑and‑play adapter to the shared expert component of Mixture‑of‑Experts models, and SEAL++, which adds an orthogonal constraint to preserve existing safety subspaces. By leveraging the always‑activated shared expert, SEAL mitigates the structural vulnerability of sparse routing to adversarial manipulation, reducing attack success rates by up to 60% with minimal impact on model capability. The approach is evaluated across six attack scenarios involving harmful prompting, jailbreaks, malicious fine‑tuning, and neuron pruning.

By Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen
arXiv AI
Aug 18

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

arXiv:2608. 15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive.

By Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li