arXiv AI

Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models

arXiv:2607. 13647v1 Announce Type: cross Abstract: Do vision models see colors the way humans do?

arXiv Computer Vision
3d ago

Perceptual Color Difference Modeling Using Machine Learning and Human Similarity Judgments

The paper presents a data‑driven method for estimating perceptual color differences by training regression models on human similarity judgments of 2,000 color pairs. Using COLIBRI fuzzy linguistic categories as features, linear regression achieves an R² of 0.595, outperforming RGB and HSI representations. The best result, an R² of 0.703, is obtained with LightGBM on a combined representation, showing that graded perceptual categories improve color‑difference prediction.

By Elnara Kadyrgali, Muragul Muratbekova, Adilet Yerkin, Nuray Toganas, Ayan Igali, Malika Ziyada, Aruzhan Burambekova, Jamaladdin Hasanov, Pakizar Shamoi
Hugging Face Trending Papers
Sep 8

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

The paper investigates how vision encoders and Vision‑Language Models (VLMs) encode conceptual information by using canonical color as a test case. By creating a dataset of objects with canonical colors and probing encoders with both color and grayscale images, the authors show that canonical color can still be decoded from grayscale inputs and is linked to predicted object identity. They further demonstrate that post‑training of VLMs can significantly influence color decodability within the vision encoder, suggesting that canonical color is a useful tool for tracing conceptual semantics in these models.

arXiv AI
Aug 13

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

arXiv:2608. 11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization.

By Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu
arXiv Computer Vision
Sep 11

HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRA

HALDETECT is a system developed for the English hallucination-detection track of ImageEval 2026, where the task is to identify the single visually grounded statement among three culturally plausible options. The approach treats the problem as a contrastive decision, outputs the answer before an explanation, and bases reasoning on colour/texture, shape/form, and context. The best model fine‑tunes Qwen2.5‑VL‑7B‑Instruct with 4‑bit QLoRA, freezes the vision encoder, and achieves a Contrastive Instability score of 0.035 on the test set, placing third among eight teams.

By Syed Mohaiminul Hoque, Md Sakhawat Hossain
arXiv AI
Sep 18

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Paint-Anything introduces a unified hex-prompt interface that allows users to specify any 24‑bit hex color for both image generation and editing. The method trains on a new Paint‑500K dataset created from real images with object grounding, perceptual color labeling, and editing‑pair synthesis, and supplements this with pure‑color anchors to address shadow‑induced color inaccuracies. Evaluated on the newly proposed Any Color Benchmark (ACBench), Paint‑Anything achieves significant improvements over the base FLUX.2‑4B model, boosting T2I and editing scores by 85.3 % and 28.3 % respectively, and outperforms competing methods on the CompColor metric.

By Ji Xie, Dewei Zhou, Xinyu Huang, Zhennan Chen, Xun Wang
arXiv AI
Aug 20

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

The paper introduces the Graded Color Attribution (GCA) dataset, a benchmark that tests whether Vision‑Language Models (VLMs) and humans can articulate and follow a threshold rule for labeling objects by color. In experiments, humans consistently adhere to their stated rules, while VLMs—despite accurately estimating color coverage—often violate their own introspective rules, especially when world‑knowledge priors are present. This discrepancy highlights a miscalibration in VLM self‑knowledge that differs from human cognition.

By Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
arXiv AI
Aug 18

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

arXiv:2608. 15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors.

By Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
arXiv Computer Vision
Aug 24

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt

The paper introduces Embedded Stroop, a diagnostic test that embeds text prompts directly into images to study interference in multimodal large language models (MLLMs). Using the What-Color-Is-the-Text (WCIT) benchmark, which tests 59 fine‑grained colors in standard, flipped, and masked conditions, the authors evaluate 16 models and find that while exact color accuracy is low (6.3%), models still recognize coarse color families (38.4%) but frequently hallucinate the embedded word instead of the true color (Stroop Hallucination Rate of 21.6%). Masking or flipping the embedded text reduces hallucinations, indicating that semantic legibility can dominate visual color perception in MLLMs.

By Jinkun Zhao, Lei Huang, Haixin Ge, Wenjun Wu