Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases).
arXiv:2606. 04409v1 Announce Type: cross Abstract: Modern deep neural networks usually have large parameter scales and nonlinear hierarchical structures, and they have achieved strong performance in computer vision.
arXiv:2602.01718v2 Announce Type: replace Abstract: Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benc...
arXiv:2607. 00371v1 Announce Type: cross Abstract: Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation.
arXiv:2608.23182v1 Announce Type: cross Abstract: We present a comparative study of label-free metrics for assessing the quality of representations in deep neural networks to understand their reliabi...
arXiv:2505. 03201v4 Announce Type: replace-cross Abstract: Integrated Gradients (IG) is a widely used attribution method in explainable AI, particularly in computer vision applications where reliable feature attribution is essential.