Hugging Face Trending Papers

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

The paper investigates how vision encoders and Vision‑Language Models (VLMs) encode conceptual information by using canonical color as a test case. By creating a dataset of objects with canonical colors and probing encoders with both color and grayscale images, the authors show that canonical color can still be decoded from grayscale inputs and is linked to predicted object identity. They further demonstrate that post‑training of VLMs can significantly influence color decodability within the vision encoder, suggesting that canonical color is a useful tool for tracing conceptual semantics in these models.

arXiv AI
Aug 10

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.

By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
arXiv AI
Aug 25

Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework

arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evide...

By Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
arXiv AI
Aug 20

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

The paper introduces the Graded Color Attribution (GCA) dataset, a benchmark that tests whether Vision‑Language Models (VLMs) and humans can articulate and follow a threshold rule for labeling objects by color. In experiments, humans consistently adhere to their stated rules, while VLMs—despite accurately estimating color coverage—often violate their own introspective rules, especially when world‑knowledge priors are present. This discrepancy highlights a miscalibration in VLM self‑knowledge that differs from human cognition.

By Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
arXiv Computer Vision
Aug 24

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt

The paper introduces Embedded Stroop, a diagnostic test that embeds text prompts directly into images to study interference in multimodal large language models (MLLMs). Using the What-Color-Is-the-Text (WCIT) benchmark, which tests 59 fine‑grained colors in standard, flipped, and masked conditions, the authors evaluate 16 models and find that while exact color accuracy is low (6.3%), models still recognize coarse color families (38.4%) but frequently hallucinate the embedded word instead of the true color (Stroop Hallucination Rate of 21.6%). Masking or flipping the embedded text reduces hallucinations, indicating that semantic legibility can dominate visual color perception in MLLMs.

By Jinkun Zhao, Lei Huang, Haixin Ge, Wenjun Wu