arXiv AI

lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation

arXiv:2606. 00022v1 Announce Type: cross Abstract: Humor generation remains difficult not only because producing fluent, novel jokes is hard, but because "funny" is audience-dependent and supervision is noisy -- preferences vary with audience, context, and culture, and annotator agreement is often low.

arXiv Computation and Language
Sep 15

IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English

arXiv:2609.15618v1 Announce Type: cross Abstract: Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving firs...

By Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol
arXiv Computation and Language
Sep 11

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

The paper introduces IRS (Incongruity-Resolution Supervision), a framework that breaks humor understanding into three parts: identifying mismatches in a visual scene, creating coherent reinterpretations of those mismatches, and aligning these interpretations with human preferences. IRS uses structured reasoning traces to guide models from visual perception to humorous interpretation, and it is evaluated on the New Yorker Cartoon Caption Contest. Experiments on 7B, 32B, and 72B models show that IRS improves caption matching and ranking, with the 72B model achieving 76.10% ranking accuracy—outperforming non-expert humans and all other multimodal baselines—and demonstrates transferable reasoning patterns in zero‑shot settings.

By Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem
arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi
arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte