arXiv AI

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms

The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.

arXiv Computation and Language
Aug 31

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

EvoHarmBench is a dynamic adversarial evaluation framework that simulates how users iteratively modify harmful content to evade moderation. It uses an optimization loop that evolves evasion strategies at the semantic-cluster level while maintaining human readability, and tests 229 semantic sub-clusters across five violation categories derived from 5,002 real-world adversarial samples. The study shows that even state‑of‑the‑art LLM‑based moderators can be bypassed with an 80.3% success rate after twelve iterations, highlighting significant vulnerabilities in current systems.

By Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding, Kangwei Liu, Ziwen Xu, Longtao Huang, Hui Xue, Ziqiang Zhu, Junjie Li, Haiwen Hong
arXiv Computation and Language
Sep 3

The Enforcement and Feasibility of Hate Speech Moderation

The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.

By Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger
arXiv Computation and Language
Sep 3

Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity

The study investigates how the demographic makeup of crowdsourced moderators influences who is protected from perceived toxic content online. Using data from 16,221 U.S. respondents who evaluated over 100,000 comments from Twitter, Reddit, and 4chan, the authors find that moderators tend to protect users who share their own demographic identities, a pattern that is amplified when moderator pools mirror the demographics of Prolific participants. Even fully representative moderator groups fail to provide equal protection, leaving Black and LGB users underprotected unless they are overrepresented.

By Zhaodi Chen, Byungkyu Lee
arXiv AI
Jul 1

How Human Feedback Shapes AI-generated Community Notes

arXiv:2606. 30905v1 Announce Type: cross Abstract: Community Notes, a bridging-based crowd-sourced fact-checking system, has emerged as a new mechanism for moderating misleading information on social media and has been adopted by major platforms including X, Facebook, Instagram, Threads, and TikTok.

By Soham De, Isaac Slaughter, Jiawei Guo, Qiao-Yun Cheng, Jiayuan Yan, Sruti Banerjee, Martin Saveski
arXiv AI
Aug 25

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

The paper presents a scalable approach to harmful content moderation on social media by leveraging large language models (LLMs) for few-shot, in-context learning. Experiments across multiple LLMs show that this method outperforms proprietary baselines such as Perspective and OpenAI Moderation, as well as prior few-shot learning techniques, in detecting harmful content. The study also explores the addition of visual cues like video thumbnails to assess multimodal improvements, highlighting the advantages of LLM-based moderation for dynamic and large-scale content filtering.

By Akash Bonagiri, Lucen Li, Rajvardhan Oak, Zeerak Babar, Magdalena Wojcieszak, Anshuman Chhabra