arXiv Computation and Language By Junyu Lu, Kaiyuan Liu, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

Read the original on arXiv Computation and Language →

The paper investigates how annotator-style rebuttals can manipulate large language model (LLM) moderation systems, either by whitewashing hateful content as normal or smearing normal content as hateful. Using a rejudge protocol that adds decision‑boundary perturbations and adversarial rationales, the authors show that such rebuttals significantly degrade moderation performance, especially in multi‑turn settings. The study finds consistent, model‑specific asymmetries between the two manipulation directions and demonstrates that explicit reasoning prompts and defensive instructions mitigate but do not eliminate the vulnerability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
3d ago

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

The paper presents an instruction‑tuned large language model (LLM) based on Qwen3 that is fine‑tuned for hate speech mitigation by unifying 36 English hate speech datasets. The authors show that this generalist LLM achieves state‑of‑the‑art performance on in‑domain benchmarks and delivers significant gains in cross‑domain and cross‑lingual generalization, outperforming specialist encoder‑based classifiers.

By Lukas Edman, Daryna Dementieva, Alexander Fraser