The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.
By Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger
arXiv:2601. 04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation.
By Lionel Z. Wang, Yusheng Zhao, Jiabin Luo, Xinfeng Li, Lixu Wang, Yinan Peng, Haoyang Li, XiaoFeng Wang, Wei Dong
arXiv:2608.29624v1 Announce Type: new
Abstract: Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text p...
By Stephen Meisenbacher, Andreea-Elena Bodea, Ahmet Bilal Ak{\i}n, Alexandra Klymenko, Jana Diesner, Florian Matthes
The paper introduces a style-aware paraphrasing method for text anonymization that leverages pretrained large language models to build compact stylistic profiles from minimal samples and rewrite text to suppress identifiable style markers while preserving meaning. It demonstrates that this approach reduces authorship attribution F1 scores by 60‑70% on blog and review datasets, outperforming both differential privacy‑based and non‑DP baselines, and maintains content quality and readability.
By Ahmed Sohair Khan, Estrid He, Monica Wachowicz, Elham Naghizade
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains...
arXiv:2608.22161v1 Announce Type: new
Abstract: Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existi...
By Qian Ma, Anna Squicciarini, Sarah Rajtmajer