arXiv Computation and Language By Ruoxuan Li, Pinqiao Wang, Sheng Li, Cameron Robert Jones

You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals

Read the original on arXiv Computation and Language →

The paper introduces a pragmatics-inspired taxonomy for evaluating how large language models (LLMs) refuse unsafe or inappropriate requests. By applying this framework to 16 modern LLMs across 14 harm categories, the authors find that while refusals are generally explicit and morally charged, they often lack interpersonal facework and instead offer safer alternatives, which can be problematic in sensitive contexts. The study argues for alignment evaluations that assess not just whether LLMs refuse, but how they do so in a contextually adaptive and socially responsible manner.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 28

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

DeflectBench is a new benchmark that evaluates how large language models (LLMs) generate rhetorical fallacies when prompted. The study tests 23,990 generations from four leading models using three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims across four controversy levels. Results show that refusal to produce fallacies depends mainly on request structure, with prompt framing and fallacy type dramatically affecting compliance rates.

By Art Kanke
arXiv AI
Jul 29

Do Models Fake Alignment Without Clear Consequences?

arXiv:2607. 24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking.

By Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao