arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.
By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
arXiv:2608. 15600v1 Announce Type: new Abstract: The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text.
By Mingyu Yuan, Shengtao Wen, Lingbing Guo, Zhen Bi, Xiang Chen
The paper proposes a training‑time explainability framework that aligns model reasoning with human‑annotated rationales to improve both classification performance and interpretability for multilingual hate speech detection. It is evaluated on HateXplain (English) and BullySent (Hinglish), datasets that capture anti‑Muslim hate in culturally coded, multilingual forms. Using methods such as LIME, Integrated Gradients, Grad‑X‑Input, and attention, the study shows that gradient‑ and attention‑based regularization boosts F‑scores, enhances plausibility and faithfulness, and captures culturally specific cues for detecting implicit anti‑Muslim hate.
By Muhammad Deedahwar Mazhar Qureshi, Sannaan Khan, Muhammad Atif Qureshi, Wael Rashwan
arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.
By Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo
The paper introduces FAID, a fine‑grained adaptive framework for detecting implicit hate speech. It first classifies samples into Shallow, Targeted, or Context‑Dependent categories and then applies tailored strategies—prompt‑tuning for shallow cases, knowledge augmentation for targeted ones, and an agentic prompt‑generation system for context‑dependent posts. Experiments on four benchmark datasets show that FAID outperforms state‑of‑the‑art baselines by allocating computational effort only where needed.
By Han Wang, Yuhu Cheng, Xuesong Wang, Yi Zhu
CIBuzzBench is a new benchmark that tests cross‑lingual understanding of Chinese internet buzzwords by providing 3,001 buzzwords with English explanations, equivalents, category labels, and harmfulness annotations. The benchmark defines three tasks—Meaning Explanation, Cross‑lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection—to evaluate how well large language models translate and interpret these culturally nuanced terms. Experiments show that current state‑of‑the‑art LLMs still struggle with fine‑grained non‑literal meanings, robust matching, and accurate harmfulness detection across languages.
By Yifan Wang, Junyu Lu, Qifan Wang, Shun Zhang, Chaozhuo Li, Jiahao Liu, Zhijun Cao, Lingbin Bu, Fanliang Bu