arXiv Computation and Language

The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI

The article examines how conversational AI systems handle requests related to intimate‑partner communication, focusing on the refusal logic that serves as a governance threshold for gendered harm. A three‑stage audit of six widely used AI models tested 1,600 prompts, identified relational framing through 300 matched pairs, and compared pre‑submission framing with post‑output critique. Findings show that most systems refused fewer than 1% of prompts, but ChatGPT 5.2 and Claude Sonnet 4.5 refused most requests, with residual leakage concentrated under intimate framing; switching from a non‑intimate to an intimate‑partner descriptor amplified non‑refusal rates dramatically, and post‑output critique did not persist across fresh sessions.

Hugging Face Trending Papers
Aug 11

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally.

arXiv Computation and Language
Aug 28

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

DeflectBench is a new benchmark that evaluates how large language models (LLMs) generate rhetorical fallacies when prompted. The study tests 23,990 generations from four leading models using three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims across four controversy levels. Results show that refusal to produce fallacies depends mainly on request structure, with prompt framing and fallacy type dramatically affecting compliance rates.

By Art Kanke
Hugging Face Trending Papers
5d ago

The Argument and the Letterhead: Source-Position Coherence in AI Evaluation

The paper investigates whether AI evaluators differentiate between an argument’s content and the source attributed to it. Using 2,976 evaluations of six fixed texts across various source attributions, the study finds that the perceived quality of an argument varies with its source, indicating source-position coherence. The authors also note that this pattern holds across topics and model configurations, and that some evaluators explicitly noted mismatches between source and position.

arXiv AI
Jun 2

VET: A Framework for Analyzing AI Discourse

arXiv:2606. 01929v1 Announce Type: new Abstract: Public discourse on AI has become polarized; exaggerated positions on AI in traditional and social media threaten the development of AI Literacy among the general public.

By Meredith Ringel Morris