arXiv AI

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

arXiv:2607. 19011v1 Announce Type: cross Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description.

arXiv Computation and Language
Sep 11

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

The paper introduces IRS (Incongruity-Resolution Supervision), a framework that breaks humor understanding into three parts: identifying mismatches in a visual scene, creating coherent reinterpretations of those mismatches, and aligning these interpretations with human preferences. IRS uses structured reasoning traces to guide models from visual perception to humorous interpretation, and it is evaluated on the New Yorker Cartoon Caption Contest. Experiments on 7B, 32B, and 72B models show that IRS improves caption matching and ranking, with the 72B model achieving 76.10% ranking accuracy—outperforming non-expert humans and all other multimodal baselines—and demonstrates transferable reasoning patterns in zero‑shot settings.

By Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem
arXiv Computation and Language
Sep 11

MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

MultiHuSE is a multimodal dataset featuring 2,407 high‑definition videos of 50 diverse actors delivering 1,463 text samples in four psychological humour styles—affiliative, aggressive, self‑enhancing, and self‑deprecating—plus neutral content. Each text is performed by multiple actors, allowing analysis of expressive diversity, and a subset includes emotion annotations. Baseline experiments show that multimodal fusion improves humour style classification accuracy over unimodal approaches, especially for affiliative humour.

By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
arXiv AI
Aug 28

TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

TransMeme introduces a multi‑agent framework for cross‑cultural meme transcreation, addressing the unique challenges of preserving intent, adapting cultural meaning, and maintaining multimodal consistency. The system coordinates specialized agents for cultural adaptation, text rewriting, revision, and visual adjustment, and is evaluated on Chinese‑English meme pairs. Human and LLM‑based evaluations show that TransMeme outperforms baselines, achieving a 33.1% average improvement in human scores and a 60% Top‑1 ranking rate in LLM judgments.

By Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He