arXiv AI By Fan Huang

Simulating Hate Speech Cascades with Multi-LLM Agents: Empirical Grounding, Modeling Fidelity, and Intervention Strategies

Read the original on arXiv AI →

arXiv:2606. 18264v1 Announce Type: cross Abstract: Faithful modeling of hateful content propagation on online platforms remains an open problem for moderation research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 24

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

The study introduces PV‑SST, a peer‑voted social‑platform testbed, and conducts a preregistered matched‑exposure experiment across four topics, four seeds, four model families, and three larger variants, totaling 448 trials. Results show that feeding agents a ranked list of prior peer posts increases lexical similarity in both the core panel and larger variants, but does not produce a reliable advantage in opinion alignment or survival rates. The only robust finding is lexical convergence driven by the peer‑ranked feed, with no consistent coordination benefit across models or topics.

By Rana Muhammad Usman, Dominic Williamson
arXiv Machine Learning
Sep 17

Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric

The paper investigates how a minority of biased agents in a multi‑agent system of large language models (LLMs) can amplify bias through textual interactions. Even a small percentage of persistently extreme agents causes significant opinion shifts among the non‑biased agents, with the effect occurring faster in the Llama 3.2 model than in a classical Friedkin‑Johnsen model. Semantic analysis shows that rhetorical consistency rises with biased exposure and that non‑biased agents adopt the biased vocabulary even when their numerical opinions change only modestly.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun
arXiv Computation and Language
Sep 3

The Enforcement and Feasibility of Hate Speech Moderation

The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.

By Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger
arXiv Computation and Language
Aug 25

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

The paper investigates how annotator-style rebuttals can manipulate large language model (LLM) moderation systems, either by whitewashing hateful content as normal or smearing normal content as hateful. Using a rejudge protocol that adds decision‑boundary perturbations and adversarial rationales, the authors show that such rebuttals significantly degrade moderation performance, especially in multi‑turn settings. The study finds consistent, model‑specific asymmetries between the two manipulation directions and demonstrates that explicit reasoning prompts and defensive instructions mitigate but do not eliminate the vulnerability.

By Junyu Lu, Kaiyuan Liu, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin