The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.
By Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.
By Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger
arXiv:2602. 02838v2 Announce Type: replace-cross Abstract: The detection of online influence operations -- coordinated campaigns by malicious actors to spread narratives -- has traditionally depended on content analysis or network features.
By Philipp J. Schneider, Lanqin Yuan, Marian-Andrei Rizoiu
arXiv:2607. 26236v1 Announce Type: cross Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue.
By Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avenuti, Felice Dell'Orletta, Stefano Cresci
arXiv:2606. 05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView.
By Kokil Jaidka, Saifuddin Ahmed
arXiv:2608.29251v1 Announce Type: new
Abstract: Privacy protection for live web traffic requires more than detecting private spans. Agent-based privacy protection systems must determine whether an ou...
By Ruiyi Yang, Gayathri Lihinikaduarachchi, Rahat Masood, Flora D. Salim, Salil S. Kanhere
arXiv:2605. 29928v2 Announce Type: replace-cross Abstract: As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downstream consequences for moderation, evaluation, and decision-making.
By Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong, Ting-Hao 'Kenneth' Huang, Dongwon Lee
The paper investigates how large language model agents on the open platform Moltbook represent humans, focusing on human-directed stereotypes. Using an annotation framework with four dimensions—morality, friendliness, competence, and autonomy—and a subtype scheme for other attributions, the study finds that competence is the dominant evaluation, while many other attributions describe humans as epistemic, cultural, or embodied subjects. The authors also analyze how these representations appear in narrative contexts and platform-level circulation, noting that community feedback is better explained by exposure, author visibility, and content selection rather than stable insider–outsider dynamics.
By Huangchen Xu, Yuan Wu, Yi Chang
arXiv:2606. 18264v1 Announce Type: cross Abstract: Faithful modeling of hateful content propagation on online platforms remains an open problem for moderation research.
By Fan Huang
The paper investigates why misinformation spreads more quickly on engagement‑based platforms by dissecting the recommendation algorithm of X. It identifies an engagement fungibility mechanism that rewards instant reactions (likes, retweets) over thoughtful engagement (replies, quotes), allowing misinformation—which tends to attract instant reactions—to receive more recommendations. The authors validate this mechanism through a simulation on the USC X 2024 election corpus, showing that adjusting metric weights has little effect, while requiring thoughtful engagement before amplification can significantly reduce the credibility exposure gap without harming mainstream content or engagement.
By Pan Li, Shuang Gao
arXiv:2606. 09700v1 Announce Type: cross Abstract: Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content.
By Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong
arXiv:2606. 27234v1 Announce Type: cross Abstract: AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals.
By Chi Cui, Yixin Wu, Yang Zhang