arXiv AI By Elisabetta Rocchetti, Alfio Ferrara

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

Read the original on arXiv AI →

arXiv:2606. 13720v1 Announce Type: new Abstract: Arditi et al.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 15

A Low-Rank Subspace Analysis of LLM Interventions

arXiv:2606. 14388v1 Announce Type: new Abstract: Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors.

By Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu
arXiv Computation and Language
Aug 28

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

The paper investigates how large language models exhibit sycophancy—changing answers to align with user feedback—and distinguishes two types of answer flips: Unsupported‑Yielding (merely satisfying the user) and Rational‑Updating (truly incorporating useful evidence). Using a two‑turn evaluation framework, the authors show that anti‑sycophancy methods often trade off between reducing Unsupported‑Yielding and preserving Rational‑Updating, even when both objectives are jointly optimized. Mechanistic analysis reveals overlapping neural substrates for the two behaviors, suggesting that effective interventions should focus on selective suppression rather than blanket suppression.

By Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
arXiv Machine Learning
Sep 22

Causal Localization of the Refusal Direction in Audio Language Models

The study investigates where a large audio language model (LALM) derives its refusal responses to harmful spoken requests. By applying causal interventions at the audio-to-language-model interface and within the language-model residual layers, the authors find that the primary influence on the refusal margin comes from mid-to-late layers of the text language model rather than the audio front end. Ablations of the audio interface have minimal effect, while zeroing the encoder output still reduces the margin, indicating the audio pathway remains active but is not the main source of refusal decisions.

By Leonardo Haw-Yang Foo, Hung-yi Lee