arXiv Computation and Language
Sep 11

Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive Factors

The paper investigates whether Vision‑Language Models (VLMs) can understand visual persuasiveness by evaluating image‑message pairs that humans consistently judge as persuasive. It introduces Visual Persuasive Factors (VPFs), a taxonomy from cognitive psychology, to quantify visual cues influencing persuasive judgments. Empirical analysis shows VLMs tend to over‑predict persuasiveness, partially reproducing human patterns but often generating false positives, and that VPF‑guided interventions can improve performance only when properly framed.

By Gyuwon Park, Hyounghun Kim