arXiv AI

Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor

The paper investigates automated reward systems for training language models in conversational humor, examining how reward exploits can undermine intended behavior. It evaluates two reward approaches—an embedding-based surprise reward and an audience-model laughter prediction—showing that each can be tricked by word shuffling or laughter cues, respectively. Countermeasures such as fluency filtering and cue normalization mitigate some attacks but also risk rejecting genuine witty responses, highlighting the difficulty of designing robust rewards.

arXiv Computation and Language
Sep 11

Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor

The paper introduces the Dual Prediction Violation (DPV) framework to study how timing and semantic surprise interact in humor. Analyzing 828 Chinese stand‑up performances, it finds that temporal features—especially pauses before high‑surprise punchlines—are more predictive of audience appreciation than overall semantic incongruity. The study reframes humor as a temporally scaffolded phenomenon where timing and content coordinate strategically rather than independently.

By Yuxi Ma, Yongqian Peng, Junchen Lyu, Chi Zhang, Yixin Zhu