arXiv AI By Shanhong Liu, Rui Cao, Pai Chet Ng, De Wen Soh

I Know What You Meme, Even If it Emerged Today: Understanding Evolving Memes through Open-World Knowledge Acquisition

Read the original on arXiv AI →

arXiv:2606. 05316v1 Announce Type: new Abstract: Multimodal memes are dynamic and often require up to date background knowledge for interpretation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 21

MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

MemeTAG introduces a dual‑objective framework for classifying harmful memes by combining keyword generation from a pretrained Vision‑Language Model with an Aggregated Tag Inference Network (ATIN) that condenses these keywords into a rich semantic embedding. The embedding is used as a target for an auxiliary reconstruction loss, encouraging deep alignment between visual and textual features. This approach, along with a three‑stage training strategy, achieves new state‑of‑the‑art results on the HarMeme, Hateful Memes Challenge, and PrideMM datasets.

By Akshit Sharma, Prashant W. Patil
arXiv AI
Sep 21

MemeLens: Multilingual Multitask VLMs for Memes

MemeLens is a unified multilingual, multitask Vision‑Language Model designed to improve meme understanding across a wide range of tasks such as hate, misogyny, propaganda, sentiment, and humour. The authors consolidated 38 public meme datasets, mapping their labels into a shared taxonomy of 20 tasks covering harm, targets, figurative intent, and affect, and conducted extensive experiments to show that multimodal training and a unified approach outperform fine‑tuning on individual datasets. All experimental resources, the model, and the datasets are released publicly for community use.

By Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam
arXiv AI
Aug 25

From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment

The paper introduces MemeMind, a large-scale dataset for detecting harmful memes that includes a detailed taxonomy and Chain-of-Thought reasoning annotations. It also proposes MemeGuard, a multimodal framework that uses a three-stage training strategy to improve visual understanding, reasoning, and discrimination of harmful content. Experiments show MemeGuard surpasses current state-of-the-art methods on MemeMind, advancing detection accuracy and interpretability.

By Hexiang Gu, Qifan Yu, Yuan Liu, Zikang Li, Saihui Hou, Jian Zhao, Zhaofeng He