arXiv:2606. 13256v1 Announce Type: cross Abstract: Humor plays a central role in human social relationships, and recent advances in computational humor create new opportunities for integrating humor into human-robot interaction (HRI).
By Anna-Maria Velentza, Anne-Gwenn Bosser
The paper introduces the Dual Prediction Violation (DPV) framework to study how timing and semantic surprise interact in humor. Analyzing 828 Chinese stand‑up performances, it finds that temporal features—especially pauses before high‑surprise punchlines—are more predictive of audience appreciation than overall semantic incongruity. The study reframes humor as a temporally scaffolded phenomenon where timing and content coordinate strategically rather than independently.
By Yuxi Ma, Yongqian Peng, Junchen Lyu, Chi Zhang, Yixin Zhu
MultiHuSE is a multimodal dataset featuring 2,407 high‑definition videos of 50 diverse actors delivering 1,463 text samples in four psychological humour styles—affiliative, aggressive, self‑enhancing, and self‑deprecating—plus neutral content. Each text is performed by multiple actors, allowing analysis of expressive diversity, and a subset includes emotion annotations. Baseline experiments show that multimodal fusion improves humour style classification accuracy over unimodal approaches, especially for affiliative humour.
By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
The paper investigates automated reward systems for training language models in conversational humor, examining how reward exploits can undermine intended behavior. It evaluates two reward approaches—an embedding-based surprise reward and an audience-model laughter prediction—showing that each can be tricked by word shuffling or laughter cues, respectively. Countermeasures such as fluency filtering and cue normalization mitigate some attacks but also risk rejecting genuine witty responses, highlighting the difficulty of designing robust rewards.
By Sam Larson
arXiv:2607. 19011v1 Announce Type: cross Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description.
By Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.
By Fabio De Ponte
arXiv:2607. 13189v1 Announce Type: cross Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese).
By Marek \v{S}uppa, Vikt\'oria Ondrejov\'a, Lucia Ganajov\'a, Gregor Karetka, Daniel Skala
arXiv:2606. 00046v1 Announce Type: cross Abstract: Video platforms such as YouTube have reshaped how users engage with entertainment and information, emphasizing brief, highly engaging content such as Shorts.
By Sydney Johns, Sanjeev Parthasarathy, Shantnu Bhalla, Vaibhav Garg
arXiv:2601.03103v2 Announce Type: replace-cross
Abstract: Humor preferences vary widely across individuals and cultures, complicating the evaluation of humor using large language models (LLMs). In th...
By Soichiro Murakami, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura
arXiv:2608.23172v1 Announce Type: new
Abstract: Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humo...
By Abhilash Nandy, Rahul Seetharaman, Aman Bansal, Rounak Saha, Manav Nitin Kapadnis, Millon Madhur Das, Pawan Goyal, Niloy Ganguly
arXiv:2607. 15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist.
By Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis
The paper introduces IRS (Incongruity-Resolution Supervision), a framework that breaks humor understanding into three parts: identifying mismatches in a visual scene, creating coherent reinterpretations of those mismatches, and aligning these interpretations with human preferences. IRS uses structured reasoning traces to guide models from visual perception to humorous interpretation, and it is evaluated on the New Yorker Cartoon Caption Contest. Experiments on 7B, 32B, and 72B models show that IRS improves caption matching and ranking, with the 72B model achieving 76.10% ranking accuracy—outperforming non-expert humans and all other multimodal baselines—and demonstrates transferable reasoning patterns in zero‑shot settings.
By Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem