arXiv Computation and Language

You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals

The paper introduces a pragmatics-inspired taxonomy for evaluating how large language models (LLMs) refuse unsafe or inappropriate requests. By applying this framework to 16 modern LLMs across 14 harm categories, the authors find that while refusals are generally explicit and morally charged, they often lack interpersonal facework and instead offer safer alternatives, which can be problematic in sensitive contexts. The study argues for alignment evaluations that assess not just whether LLMs refuse, but how they do so in a contextually adaptive and socially responsible manner.

arXiv Computation and Language
Aug 28

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

DeflectBench is a new benchmark that evaluates how large language models (LLMs) generate rhetorical fallacies when prompted. The study tests 23,990 generations from four leading models using three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims across four controversy levels. Results show that refusal to produce fallacies depends mainly on request structure, with prompt framing and fallacy type dramatically affecting compliance rates.

By Art Kanke
arXiv AI
Jul 29

Do Models Fake Alignment Without Clear Consequences?

arXiv:2607. 24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking.

By Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao
arXiv AI
Jul 23

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

arXiv:2607. 19629v1 Announce Type: cross Abstract: Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other.

By Eunna Lee
arXiv Computation and Language
Aug 25

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

The paper investigates how annotator-style rebuttals can manipulate large language model (LLM) moderation systems, either by whitewashing hateful content as normal or smearing normal content as hateful. Using a rejudge protocol that adds decision‑boundary perturbations and adversarial rationales, the authors show that such rebuttals significantly degrade moderation performance, especially in multi‑turn settings. The study finds consistent, model‑specific asymmetries between the two manipulation directions and demonstrates that explicit reasoning prompts and defensive instructions mitigate but do not eliminate the vulnerability.

By Junyu Lu, Kaiyuan Liu, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin