arXiv Computation and Language
Sep 11

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

The paper introduces IRS (Incongruity-Resolution Supervision), a framework that breaks humor understanding into three parts: identifying mismatches in a visual scene, creating coherent reinterpretations of those mismatches, and aligning these interpretations with human preferences. IRS uses structured reasoning traces to guide models from visual perception to humorous interpretation, and it is evaluated on the New Yorker Cartoon Caption Contest. Experiments on 7B, 32B, and 72B models show that IRS improves caption matching and ranking, with the 72B model achieving 76.10% ranking accuracy—outperforming non-expert humans and all other multimodal baselines—and demonstrates transferable reasoning patterns in zero‑shot settings.

By Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem
arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte