arXiv Computation and Language By Zhaodi Chen, Byungkyu Lee

Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity

Read the original on arXiv Computation and Language →

The study investigates how the demographic makeup of crowdsourced moderators influences who is protected from perceived toxic content online. Using data from 16,221 U.S. respondents who evaluated over 100,000 comments from Twitter, Reddit, and 4chan, the authors find that moderators tend to protect users who share their own demographic identities, a pattern that is amplified when moderator pools mirror the demographics of Prolific participants. Even fully representative moderator groups fail to provide equal protection, leaving Black and LGB users underprotected unless they are overrepresented.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 12

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms

The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.

By Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
arXiv Computation and Language
Sep 3

The Enforcement and Feasibility of Hate Speech Moderation

The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.

By Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger