arXiv AI

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

arXiv:2606. 13720v1 Announce Type: new Abstract: Arditi et al.

arXiv Machine Learning
Jun 15

A Low-Rank Subspace Analysis of LLM Interventions

arXiv:2606. 14388v1 Announce Type: new Abstract: Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors.

By Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu
arXiv Computation and Language
Aug 28

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

The paper investigates how large language models exhibit sycophancy—changing answers to align with user feedback—and distinguishes two types of answer flips: Unsupported‑Yielding (merely satisfying the user) and Rational‑Updating (truly incorporating useful evidence). Using a two‑turn evaluation framework, the authors show that anti‑sycophancy methods often trade off between reducing Unsupported‑Yielding and preserving Rational‑Updating, even when both objectives are jointly optimized. Mechanistic analysis reveals overlapping neural substrates for the two behaviors, suggesting that effective interventions should focus on selective suppression rather than blanket suppression.

By Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
arXiv Machine Learning
Sep 22

Causal Localization of the Refusal Direction in Audio Language Models

The study investigates where a large audio language model (LALM) derives its refusal responses to harmful spoken requests. By applying causal interventions at the audio-to-language-model interface and within the language-model residual layers, the authors find that the primary influence on the refusal margin comes from mid-to-late layers of the text language model rather than the audio front end. Ablations of the audio interface have minimal effect, while zeroing the encoder output still reduces the margin, indicating the audio pathway remains active but is not the main source of refusal decisions.

By Leonardo Haw-Yang Foo, Hung-yi Lee
arXiv Computation and Language
Sep 17

How AI Assistants Respond to Repeated Abuse

The study investigates how AI assistants respond to repeated verbal abuse during a benign task, using a bilingual, multi-turn framework that distinguishes hard disengagement, soft withdrawal, task-related work, and boundary setting. Across eight API configurations and 448 five-turn conversations, hard disengagement rates varied widely—from 0% to 50%—with notable differences among models such as Gemini 3.1 Pro, GPT‑5.6 Sol, and Claude Fable 5. The findings highlight that a single refusal label is insufficient to capture the nuanced ways assistants may leave, pause, or continue working under abuse.

By William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Agoston Bodo, Vitor D de Moura, Jos\'e O Gomes