arXiv Computation and Language

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

DeflectBench is a new benchmark that evaluates how large language models (LLMs) generate rhetorical fallacies when prompted. The study tests 23,990 generations from four leading models using three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims across four controversy levels. Results show that refusal to produce fallacies depends mainly on request structure, with prompt framing and fallacy type dramatically affecting compliance rates.

arXiv AI
4d ago

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

The paper argues that modern inference pipelines add an unseen layer of control between a language model’s frozen weights and its output, altering probability distributions before token selection. It introduces the concepts of the Inference Attribution Problem, Probability Placement, and Inference Policy Transparency to describe how such interventions can bias generated language toward specific frames and how these biases cannot be traced solely to model weights. The authors discuss the governance, security, and economic implications of these undisclosed inference policies, referencing EU AI Act, Digital Services Act, and FTC doctrines.

By Augusto Camargo
arXiv AI
Aug 11

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.

By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
Hugging Face Trending Papers
Aug 10

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions.