arXiv Machine Learning
Jun 4

Expert-Aware Refusal Steering

arXiv:2606. 04160v1 Announce Type: cross Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed requests.

By Anna C. Marbut, Daniel R. Olson, Travis J. Wheeler
arXiv AI
1d ago

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

The study investigates whether the door‑in‑the‑face technique—making a large request that is refused to increase the likelihood of a smaller follow‑up request being granted—works on large language models. Nine production models from Anthropic, OpenAI, Google, and Haiku were tested; the technique succeeded on Anthropic’s frontier models but backfired on the others. The effect depends on the model family and the content of the request, and it does not transfer to refusals from public benchmarks.

By Til Jordan
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.