arXiv:2607. 13647v1 Announce Type: cross Abstract: Do vision models see colors the way humans do?
By Ayan Igali, Pakizar Shamoi
arXiv:2609.14495v1 Announce Type: new
Abstract: Image colorization is an inherently ill-posed task, since a single grayscale image may correspond to multiple plausible colorized results. Consequently...
By Yunkai Zhuang, Qihang Yan, Zicheng Zhang, Guangtao Zhai
arXiv:2608. 14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation.
By Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
arXiv:2609.09124v1 Announce Type: cross
Abstract: Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, inf...
By Xiaofu Chen, Stella Frank, Yova Kementchedjhieva
Paint-Anything introduces a unified hex-prompt interface that allows users to specify any 24‑bit hex color for both image generation and editing. The method trains on a new Paint‑500K dataset created from real images with object grounding, perceptual color labeling, and editing‑pair synthesis, and supplements this with pure‑color anchors to address shadow‑induced color inaccuracies. Evaluated on the newly proposed Any Color Benchmark (ACBench), Paint‑Anything achieves significant improvements over the base FLUX.2‑4B model, boosting T2I and editing scores by 85.3 % and 28.3 % respectively, and outperforms competing methods on the CompColor metric.
By Ji Xie, Dewei Zhou, Xinyu Huang, Zhennan Chen, Xun Wang
The paper investigates how vision encoders and Vision‑Language Models (VLMs) encode conceptual information by using canonical color as a test case. By creating a dataset of objects with canonical colors and probing encoders with both color and grayscale images, the authors show that canonical color can still be decoded from grayscale inputs and is linked to predicted object identity. They further demonstrate that post‑training of VLMs can significantly influence color decodability within the vision encoder, suggesting that canonical color is a useful tool for tracing conceptual semantics in these models.