arXiv:2606. 05316v1 Announce Type: new Abstract: Multimodal memes are dynamic and often require up to date background knowledge for interpretation.
By Shanhong Liu, Rui Cao, Pai Chet Ng, De Wen Soh
arXiv:2609.26907v1 Announce Type: cross
Abstract: Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal cla...
By Akshit Sharma, Prashant W. Patil
The paper introduces MemeMind, a large-scale dataset for detecting harmful memes that includes a detailed taxonomy and Chain-of-Thought reasoning annotations. It also proposes MemeGuard, a multimodal framework that uses a three-stage training strategy to improve visual understanding, reasoning, and discrimination of harmful content. Experiments show MemeGuard surpasses current state-of-the-art methods on MemeMind, advancing detection accuracy and interpretability.
By Hexiang Gu, Qifan Yu, Yuan Liu, Zikang Li, Saihui Hou, Jian Zhao, Zhaofeng He
arXiv:2607. 15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist.
By Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis
MemeLens is a unified multilingual, multitask Vision‑Language Model designed to improve meme understanding across a wide range of tasks such as hate, misogyny, propaganda, sentiment, and humour. The authors consolidated 38 public meme datasets, mapping their labels into a shared taxonomy of 20 tasks covering harm, targets, figurative intent, and affect, and conducted extensive experiments to show that multimodal training and a unified approach outperform fine‑tuning on individual datasets. All experimental resources, the model, and the datasets are released publicly for community use.
By Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam
arXiv:2603. 01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.
By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi