arXiv AI By Eyal Hanania, Daniel Arkushin, Naveh Ayal, Jonathan Benvenisti, Amos Bercovich, Elie Zemmour, Sahar Froim

When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

Read the original on arXiv AI →

The study investigates structured disagreement among annotators in temporal laughter localization, revealing that disagreements are not random but exhibit systematic patterns—larger at offsets than onsets, more frequent for chuckles than full laughs, and predictable from event attributes. Re-annotating the SMILE-Temporal benchmark with multiple annotators per video shows that evaluating against a single reference annotation can significantly bias system scores and ranking accuracy. The authors propose a disagreement‑calibrated evaluation using conformally calibrated tolerance bands that better reflect the full annotator distribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor

The paper introduces the Dual Prediction Violation (DPV) framework to study how timing and semantic surprise interact in humor. Analyzing 828 Chinese stand‑up performances, it finds that temporal features—especially pauses before high‑surprise punchlines—are more predictive of audience appreciation than overall semantic incongruity. The study reframes humor as a temporally scaffolded phenomenon where timing and content coordinate strategically rather than independently.

By Yuxi Ma, Yongqian Peng, Junchen Lyu, Chi Zhang, Yixin Zhu
arXiv Computer Vision
2d ago

Predicting Human Disagreement for Calibrated Dynamic Facial Expression Recognition

The paper introduces a disagreement‑aware dynamic facial expression recognition framework that directly learns from raw annotator vote vectors using a Dirichlet‑Multinomial likelihood, preserving both predictive mean and scale‑sensitive supervision. It adds an ambiguity head to estimate annotation entropy for unseen clips and employs a Chow‑style reject rule that integrates ambiguity, vacuity, temporal instability, and input quality for selective prediction. On the DFEW benchmark, the method maintains recognition accuracy while cutting expected calibration error by 30 % and area‑under‑risk‑curve by 15 %, with predicted ambiguity correlating 0.52 (Spearman) with true annotation entropy, and these gains transfer to FERV39k and hold under identity‑ and movie‑disjoint splits.

By Yiming Wang, Frederick W. B. Li, Jingyun Wang