arXiv Computer Vision

Perceive, Refine, Reason: A Calibrated Pipeline for Measuring Indicators in Strategic Visual Communication on Social Media

arXiv Computation and Language
Sep 11

Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive Factors

The paper investigates whether Vision‑Language Models (VLMs) can understand visual persuasiveness by evaluating image‑message pairs that humans consistently judge as persuasive. It introduces Visual Persuasive Factors (VPFs), a taxonomy from cognitive psychology, to quantify visual cues influencing persuasive judgments. Empirical analysis shows VLMs tend to over‑predict persuasiveness, partially reproducing human patterns but often generating false positives, and that VPF‑guided interventions can improve performance only when properly framed.

By Gyuwon Park, Hyounghun Kim
arXiv Computer Vision
Sep 16

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

The paper introduces ViD, a vision‑dominant gender bias mitigation framework for large vision‑language models. ViD uses causal analysis of attention patterns and dual mechanisms—backdoor adjustment and refined token selection—to suppress bias while preserving reasoning and generation quality. Experiments show a 14.7% reduction in gender bias on FACET and significant improvements on MS COCO image captioning, all without extra training overhead.

By Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu, Youwei Zhao, Ruichun Tang