arXiv AI
Aug 18

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

arXiv:2602. 18094v2 Announce Type: replace-cross Abstract: Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID).

By Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen
arXiv Computer Vision
3d ago

SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

SalArt-VQA is a diagnostic benchmark that tests whether vision‑language models can understand salient artifacts in AI‑generated images. It includes 950 images and 3,681 multiple‑choice questions that assess artifact presence, semantic localization, spatial grounding, and evidence‑grounded defect identification. The benchmark reveals that high image‑level detection accuracy can mask failures in grounded understanding, showing a trade‑off between sensitivity and calibration.

By Ruotian Zhang, Xiaoxiao Sun, Junzhe Huang, James Burgess, Serena Yeung-Levy