The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE.
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.
The paper introduces ‘Fool’s Gold’, a defensive deception technique for open‑weight language models that hardens them against safety‑removal attacks. By training decoy responses within a differentiable simulation of the attack, the method poisons the payoff of stripped refusal mechanisms, producing confident but falsified answers to hazardous requests while preserving benign behavior. Experiments on seven models (9B‑122B) show that 51‑90% of attacked‑state responses become decoys, with the defense accounting for 27‑84% of this effect, and that the defended 122B model remains within benign‑behavior budgets.
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
arXiv:2607. 27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing.