Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper investigates how vision encoders and Vision‑Language Models (VLMs) encode conceptual information by using canonical color as a test case. By creating a dataset of objects with canonical colors and probing encoders with both color and grayscale images, the authors show that canonical color can still be decoded from grayscale inputs and is linked to predicted object identity. They further demonstrate that post‑training of VLMs can significantly influence color decodability within the vision encoder, suggesting that canonical color is a useful tool for tracing conceptual semantics in these models.
arXiv:2604.03114v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) may need to forget visual concepts after deployment because of privacy, copyright, licensing, safety, or policy...
arXiv:2607. 13647v1 Announce Type: cross Abstract: Do vision models see colors the way humans do?
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
arXiv:2511. 19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.
arXiv:2606. 20077v1 Announce Type: cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals.