A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 04160v1 Announce Type: cross Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed requests.
arXiv:2608.29109v1 Announce Type: new Abstract: Large language models often answer structurally unanswerable questions, such as computing cot(-540{\deg}) or evaluating (1).startswith("1"), instead of...
The study investigates whether the door‑in‑the‑face technique—making a large request that is refused to increase the likelihood of a smaller follow‑up request being granted—works on large language models. Nine production models from Anthropic, OpenAI, Google, and Haiku were tested; the technique succeeded on Anthropic’s frontier models but backfired on the others. The effect depends on the model family and the content of the request, and it does not transfer to refusals from public benchmarks.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2605. 21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations.