arXiv Computation and Language By Art Kanke

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

Read the original on arXiv Computation and Language →

DeflectBench is a new benchmark that evaluates how large language models (LLMs) generate rhetorical fallacies when prompted. The study tests 23,990 generations from four leading models using three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims across four controversy levels. Results show that refusal to produce fallacies depends mainly on request structure, with prompt framing and fallacy type dramatically affecting compliance rates.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
4d ago

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

The paper argues that modern inference pipelines add an unseen layer of control between a language model’s frozen weights and its output, altering probability distributions before token selection. It introduces the concepts of the Inference Attribution Problem, Probability Placement, and Inference Policy Transparency to describe how such interventions can bias generated language toward specific frames and how these biases cannot be traced solely to model weights. The authors discuss the governance, security, and economic implications of these undisclosed inference policies, referencing EU AI Act, Digital Services Act, and FTC doctrines.

By Augusto Camargo