Where a Model Sends Its Own Repeated Token
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
The paper demonstrates that a single-direction white‑box attack, which projects a ‘refusal direction’ from a language model’s weights, remains effective against a 320B‑parameter mixture‑of‑experts (MoE) model (GLM‑5.3‑Flash). The attack requires only a few hundred contrastive prompts and no gradient training, and it reduces refusal behavior by up to 89 percentage points across seven harmful‑content benchmarks while leaving overall capability unchanged. However, the attack’s impact is distributed across multiple components—attention, dense, and routed‑expert writers—so that only a joint intervention removes most of the refusal ability, and the conventional module‑name matching approach fails to capture this effect in MoE architectures.
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
The paper introduces AMRA, a weight‑editing technique that mitigates abliteration—an attack that removes refusal capabilities from large language models by projecting weight matrices orthogonal to a refusal direction. AMRA obscures the refusal signal through rank‑$k$ updates to residual stream writer matrices, replaces refusal‑inducing activations with random aliases, and adjusts downstream reader matrices to maintain original behavior. Experiments on Llama‑3‑8B and Gemma‑2‑9B show significant improvements in post‑abliteration refusal scores with minimal impact on overall model performance.
arXiv:2609.06934v1 Announce Type: cross Abstract: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi...
arXiv:2608. 11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters.
arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.
The paper introduces "capability laundering," a method where a weaker, unaligned language model splits a harmful task into benign subproblems, consults a stronger aligned model on each, and locally combines the answers. Experiments with GPT‑5.5, Claude Opus 4.8, and Grok‑4.3 as consultants to various local orchestrators show significant uplift on CyBench, BountyBench, and a bioweapon attack chain, with Gemma‑4‑31B recovering many more candidates than the aligned models alone. The results reveal that refusing a harmful task does not stop frontier capabilities from being transferred and composed across multiple permitted interactions.
The paper investigates post‑training quantization of transformer attention blocks by optimizing a joint loss over the Q, K, V projections rather than individual weight matrices. Using this joint attention‑based objective (JAB), the authors achieve significant compression on Mistral‑7B, recovering 77‑90% of the performance gap at 3 bits, but the method fails when MLP layers are included. A role‑aware offset rule that ignores sensitivity estimates outperforms JAB on GPT‑2 and full Mistral‑7B, demonstrating that the matrix a weight belongs to is more critical than sensitivity metrics.
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
arXiv:2607.27836v2 Announce Type: replace Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recover...
arXiv:2508.20766v2 Announce Type: replace-cross Abstract: Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent researc...
arXiv:2609.16204v1 Announce Type: new Abstract: Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects...