arXiv Computation and Language By Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

Read the original on arXiv Computation and Language →

The paper introduces IRS (Incongruity-Resolution Supervision), a framework that breaks humor understanding into three parts: identifying mismatches in a visual scene, creating coherent reinterpretations of those mismatches, and aligning these interpretations with human preferences. IRS uses structured reasoning traces to guide models from visual perception to humorous interpretation, and it is evaluated on the New Yorker Cartoon Caption Contest. Experiments on 7B, 32B, and 72B models show that IRS improves caption matching and ranking, with the 72B model achieving 76.10% ranking accuracy—outperforming non-expert humans and all other multimodal baselines—and demonstrates transferable reasoning patterns in zero‑shot settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 2

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

arXiv:2606. 02578v1 Announce Type: cross Abstract: Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visual evidence conflicts with textual cues, MLLM judges tend to reward plausible narratives over perceptually correct answers.

By Seojeong Park, Jiho Choi, Junyong Kang, Seonho Lee, Jaeyo Shin, Hyunjung Shim