The paper introduces FAID, a fine‑grained adaptive framework for detecting implicit hate speech. It first classifies samples into Shallow, Targeted, or Context‑Dependent categories and then applies tailored strategies—prompt‑tuning for shallow cases, knowledge augmentation for targeted ones, and an agentic prompt‑generation system for context‑dependent posts. Experiments on four benchmark datasets show that FAID outperforms state‑of‑the‑art baselines by allocating computational effort only where needed.
By Han Wang, Yuhu Cheng, Xuesong Wang, Yi Zhu
arXiv:2605.31349v2 Announce Type: replace-cross
Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - con...
By Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand, Saurabh Kumar Pandey, Animesh Mukherjee
The paper introduces ProKDA, a progressive knowledge-to-decision alignment framework for explainable hateful meme detection. ProKDA separates explanation generation and label prediction into three sequential training stages—background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment—reducing task interference. Experiments on three public benchmarks demonstrate that ProKDA achieves state‑of‑the‑art detection performance while providing accurate, evidence‑supported explanations for moderation decisions.
By Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia
MexHat is a newly released video dataset aimed at improving hate‑speech detection in Mexican Spanish. It contains roughly 1,000 clips annotated for three broad categories—no negative content, offensive content, and hate‑speech—as well as a finer classification into three hate‑speech sub‑categories. The paper presents dataset statistics and baseline results, underscoring the challenges of detecting culturally and contextually nuanced hate speech in multimodal content.
By Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
arXiv:2009. 10277v2 Announce Type: replace-cross Abstract: We propose a system for measuring hate speech on a continuous, interval-valued spectrum ranging from genocidal to supportive speech by combining supervised deep learning with faceted Rasch item response theory (IRT).
By Chris J. Kennedy, Geoff Bacon, Alexander Sahn, Claudia von Vacano
Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels.