arXiv AI By Ayan Igali, Pakizar Shamoi

Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models

Read the original on arXiv AI →

arXiv:2607. 13647v1 Announce Type: cross Abstract: Do vision models see colors the way humans do?

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
6d ago

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

arXiv:2608. 11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization.

By Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu
arXiv AI
1d ago

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

arXiv:2608. 15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors.

By Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
arXiv AI
Jun 24

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.

By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi