arXiv AI By Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

Read the original on arXiv AI →

arXiv:2608. 14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 24

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt

The paper introduces Embedded Stroop, a diagnostic test that embeds text prompts directly into images to study interference in multimodal large language models (MLLMs). Using the What-Color-Is-the-Text (WCIT) benchmark, which tests 59 fine‑grained colors in standard, flipped, and masked conditions, the authors evaluate 16 models and find that while exact color accuracy is low (6.3%), models still recognize coarse color families (38.4%) but frequently hallucinate the embedded word instead of the true color (Stroop Hallucination Rate of 21.6%). Masking or flipping the embedded text reduces hallucinations, indicating that semantic legibility can dominate visual color perception in MLLMs.

By Jinkun Zhao, Lei Huang, Haixin Ge, Wenjun Wu
arXiv AI
Aug 20

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

The paper introduces the Graded Color Attribution (GCA) dataset, a benchmark that tests whether Vision‑Language Models (VLMs) and humans can articulate and follow a threshold rule for labeling objects by color. In experiments, humans consistently adhere to their stated rules, while VLMs—despite accurately estimating color coverage—often violate their own introspective rules, especially when world‑knowledge priors are present. This discrepancy highlights a miscalibration in VLM self‑knowledge that differs from human cognition.

By Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
arXiv Computer Vision
Sep 16

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

The paper introduces ViD, a vision‑dominant gender bias mitigation framework for large vision‑language models. ViD uses causal analysis of attention patterns and dual mechanisms—backdoor adjustment and refined token selection—to suppress bias while preserving reasoning and generation quality. Experiments show a 14.7% reduction in gender bias on FACET and significant improvements on MS COCO image captioning, all without extra training overhead.

By Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu, Youwei Zhao, Ruichun Tang