arXiv:2606. 28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies.
By Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy
arXiv:2609.06245v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often str...
By Yixin Wan, Tianle Zheng, Kai-Wei Chang
The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.
By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.
By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
The paper introduces MOSAIC, a method for personalized federated vision‑language models that addresses domain heterogeneity by focusing on class‑specific cross‑domain residuals. It constructs a decision‑aware harmfulness score to identify residuals that hurt image‑text decision margins, then uses a low‑rank residual adapter with shared class factors and private domain factors, an image‑conditioned gate, and harmful‑pair‑aware reweighting to refine local updates. Experiments on Office31, OfficeHome, and DomainNet100 show consistent improvements in macro‑client top‑1 accuracy across various domain‑shift scenarios.
By Wentao Yue, Qingyu Mao, Tianyou Lai, Ahmed M. Abdelmoniem, Qilei Li
The paper introduces Percept-V, a dataset of 6,000 program-generated images across 30 domains designed to test simple visual perception skills from the TVPS-4 framework. Experiments show that state‑of‑the‑art multimodal large language models perform poorly compared to humans, especially as image complexity increases, and that fine‑tuning yields only limited generalization to related datasets. The study highlights specific perception skills that remain challenging for current models.
By Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu
arXiv:2608.28696v1 Announce Type: new
Abstract: Visual in-context learning (ICL) with multimodal large language models (MLLMs) is effective for fine-grained visual classification, but each retrieved...
By Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury
Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases).
arXiv:2607. 18615v1 Announce Type: cross Abstract: Machine unlearning for vision-language models (VLMs) remains underexplored.
By Zijie Liu, Jinhao Duan, Gaowen Liu, Sijia Liu, Tianlong Chen
arXiv:2606. 08970v1 Announce Type: new Abstract: Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerous VLM candidates.
By Can Wang, Shengwei Wang, Bolin Zhang, Zhiying Tu, Dianhui Chu
arXiv:2609.01027v1 Announce Type: new
Abstract: Out-of-distribution (OOD) detection predicts whether a test image belongs to none of the predefined classes. To evaluate this task, benchmarks need ima...
By Ruslan Rozumnyi, Mat\v{e}j Such\'anek, Tom\'a\v{s} Voj\'i\v{r}, Kl\'ara Janou\v{s}kov\'a, Ji\v{r}\'i Matas