arXiv Computation and Language By Yi Shi, Tanyu Chen, Kai Shen

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Read the original on arXiv Computation and Language →

The paper demonstrates that a single-direction white‑box attack, which projects a ‘refusal direction’ from a language model’s weights, remains effective against a 320B‑parameter mixture‑of‑experts (MoE) model (GLM‑5.3‑Flash). The attack requires only a few hundred contrastive prompts and no gradient training, and it reduces refusal behavior by up to 89 percentage points across seven harmful‑content benchmarks while leaving overall capability unchanged. However, the attack’s impact is distributed across multiple components—attention, dense, and routed‑expert writers—so that only a joint intervention removes most of the refusal ability, and the conventional module‑name matching approach fails to capture this effect in MoE architectures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 7

Locating and Steering Refusal Beyond Attention

The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.

By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv AI
Aug 20

Abliteration Mitigation via Refusal Aliases

The paper introduces AMRA, a weight‑editing technique that mitigates abliteration—an attack that removes refusal capabilities from large language models by projecting weight matrices orthogonal to a refusal direction. AMRA obscures the refusal signal through rank‑$k$ updates to residual stream writer matrices, replaces refusal‑inducing activations with random aliases, and adjusts downstream reader matrices to maintain original behavior. Experiments on Llama‑3‑8B and Gemma‑2‑9B show significant improvements in post‑abliteration refusal scores with minimal impact on overall model performance.

By Nathan Truong