arXiv:2605. 26772v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal.
By Kia-J\"ung Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp
arXiv:2603.13359v2 Announce Type: replace
Abstract: Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusa...
By Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda
The paper challenges the notion that refusal in large language models is governed by a single direction in activation space. It demonstrates that different refusal and non‑compliance categories map to distinct geometric directions, yet steering along any of these directions yields similar refusal–over‑refusal trade‑offs, acting as a shared one‑dimensional control knob. Using sparse autoencoders, the authors reveal a structured internal representation of refusal, comprising a reusable core of shared latents and style‑ or domain‑specific latents, and show that linear interventions collapse this structure into uniform behavioral control.
By Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, Husrev Taha Sencar
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
By Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata
arXiv:2603.27518v4 Announce Type: replace
Abstract: Aligned language models that are trained to refuse harmful requests also exhibit over-refusal: they decline safe instructions that seemingly resemb...
By Utsav Maskey, Mark Dras, Usman Naseem
The paper introduces Persona Dosing, a method that uses an activation‑steering coefficient to control the intensity of a language model’s persona traits. By conditioning a FLAS controller on trait descriptions and calibrating its flow time against measured trait expression, the approach can adjust trait intensity without requiring paired training data. Experiments on Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑3‑4B show significant increases in core‑trait expression and low targeting errors across multiple traits.
By Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen
arXiv:2605. 21006v2 Announce Type: replace Abstract: We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect.
By Ishaan Kelkar, Nebras Alam, Vikram Kakaria, Madhur Panwar, Vasu Sharma, Maheep Chaudhary
arXiv:2607. 13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.
By Winston Zeng, Ali Emami, Jinho Choi
arXiv:2609.25021v1 Announce Type: new
Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such...
By J\k{e}drzej Maczan
The study investigates where a large audio language model (LALM) derives its refusal responses to harmful spoken requests. By applying causal interventions at the audio-to-language-model interface and within the language-model residual layers, the authors find that the primary influence on the refusal margin comes from mid-to-late layers of the text language model rather than the audio front end. Ablations of the audio interface have minimal effect, while zeroing the encoder output still reduces the margin, indicating the audio pathway remains active but is not the main source of refusal decisions.
By Leonardo Haw-Yang Foo, Hung-yi Lee
The paper investigates how refusal training shapes the internal geometry of language models, showing that activation updates from refusal-completion losses create a distinct low‑dimensional refusal subspace. In a case study on OLMo‑2‑0425‑1B‑Instruct, the authors link the brittleness of refusal directions to repetitive refusal prefixes and demonstrate that using diverse refusal starts can increase the stable rank of gradients, thereby hardening the model against vector‑ablation attacks. The work provides insights into the emergence of safety‑critical features and offers a potential strategy to strengthen refusal robustness.
By Andrey Labunets
arXiv:2608. 10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making.
By Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou