arXiv Computation and Language

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

The paper demonstrates that a single-direction white‑box attack, which projects a ‘refusal direction’ from a language model’s weights, remains effective against a 320B‑parameter mixture‑of‑experts (MoE) model (GLM‑5.3‑Flash). The attack requires only a few hundred contrastive prompts and no gradient training, and it reduces refusal behavior by up to 89 percentage points across seven harmful‑content benchmarks while leaving overall capability unchanged. However, the attack’s impact is distributed across multiple components—attention, dense, and routed‑expert writers—so that only a joint intervention removes most of the refusal ability, and the conventional module‑name matching approach fails to capture this effect in MoE architectures.

arXiv Machine Learning
Sep 7

Locating and Steering Refusal Beyond Attention

The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.

By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv AI
Aug 20

Abliteration Mitigation via Refusal Aliases

The paper introduces AMRA, a weight‑editing technique that mitigates abliteration—an attack that removes refusal capabilities from large language models by projecting weight matrices orthogonal to a refusal direction. AMRA obscures the refusal signal through rank‑$k$ updates to residual stream writer matrices, replaces refusal‑inducing activations with random aliases, and adjusts downstream reader matrices to maintain original behavior. Experiments on Llama‑3‑8B and Gemma‑2‑9B show significant improvements in post‑abliteration refusal scores with minimal impact on overall model performance.

By Nathan Truong
arXiv AI
Sep 15

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

The paper introduces "capability laundering," a method where a weaker, unaligned language model splits a harmful task into benign subproblems, consults a stronger aligned model on each, and locally combines the answers. Experiments with GPT‑5.5, Claude Opus 4.8, and Grok‑4.3 as consultants to various local orchestrators show significant uplift on CyBench, BountyBench, and a bioweapon attack chain, with Gemma‑4‑31B recovering many more candidates than the aligned models alone. The results reveal that refusing a harmful task does not stop frontier capabilities from being transferred and composed across multiple permitted interactions.

By Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, Ahmed Salem
Hugging Face Trending Papers
2d ago

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

The paper investigates post‑training quantization of transformer attention blocks by optimizing a joint loss over the Q, K, V projections rather than individual weight matrices. Using this joint attention‑based objective (JAB), the authors achieve significant compression on Mistral‑7B, recovering 77‑90% of the performance gap at 3 bits, but the method fails when MLP layers are included. A role‑aware offset rule that ignores sensitivity estimates outperforms JAB on GPT‑2 and full Mistral‑7B, demonstrating that the matrix a weight belongs to is more critical than sensitivity metrics.