MemeTAG introduces a dual‑objective framework for classifying harmful memes by combining keyword generation from a pretrained Vision‑Language Model with an Aggregated Tag Inference Network (ATIN) that condenses these keywords into a rich semantic embedding. The embedding is used as a target for an auxiliary reconstruction loss, encouraging deep alignment between visual and textual features. This approach, along with a three‑stage training strategy, achieves new state‑of‑the‑art results on the HarMeme, Hateful Memes Challenge, and PrideMM datasets.
By Akshit Sharma, Prashant W. Patil
MemeLens is a unified multilingual, multitask Vision‑Language Model designed to improve meme understanding across a wide range of tasks such as hate, misogyny, propaganda, sentiment, and humour. The authors consolidated 38 public meme datasets, mapping their labels into a shared taxonomy of 20 tasks covering harm, targets, figurative intent, and affect, and conducted extensive experiments to show that multimodal training and a unified approach outperform fine‑tuning on individual datasets. All experimental resources, the model, and the datasets are released publicly for community use.
By Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam
arXiv:2609.26907v1 Announce Type: cross
Abstract: Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal cla...
By Akshit Sharma, Prashant W. Patil
arXiv:2607. 03981v1 Announce Type: cross Abstract: Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah
The paper introduces MemeMind, a large-scale dataset for detecting harmful memes that includes a detailed taxonomy and Chain-of-Thought reasoning annotations. It also proposes MemeGuard, a multimodal framework that uses a three-stage training strategy to improve visual understanding, reasoning, and discrimination of harmful content. Experiments show MemeGuard surpasses current state-of-the-art methods on MemeMind, advancing detection accuracy and interpretability.
By Hexiang Gu, Qifan Yu, Yuan Liu, Zikang Li, Saihui Hou, Jian Zhao, Zhaofeng He
arXiv:2609.13794v1 Announce Type: new
Abstract: Detecting harmful memes is critical for maintaining safe online communities. However, harmful intent is often implicit, arising from visual-textual inc...
By Hanling Wang, Chenlong Wei, Yingjuan Li, Di Wu, Yuchao Zhang, Xiaohui Zhu, Yao Zhu