arXiv Computer Vision By Nikita Kisel, Illia Volkov, Klara Janouskova, Jiri Matas

Multimodal Large Language Models as Image Classifiers

Read the original on arXiv Computer Vision →

The paper investigates how evaluation protocols and ground‑truth quality affect the classification performance of Multimodal Large Language Models (MLLMs). It identifies and corrects key issues such as discarded out‑of‑list outputs, inflated distractor choices, and poor open‑world mapping, and shows that design choices like batch size, image ordering, and text encoder selection significantly influence accuracy. Using a multilabel reannotation of 625 ImageNet‑1k classes (ReGT), the study finds that corrected labels can boost MLLM performance by up to 10.8%, narrowing the gap with supervised models, and demonstrates that MLLMs can assist human annotators in about half of difficult cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Jun 30

DataComp-VLM: Improved Open Datasets for Vision-Language Models

arXiv:2606. 28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies.

By Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy
arXiv Computer Vision
3d ago

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.

By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
arXiv AI
Sep 10

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.

By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
arXiv Computer Vision
Oct 2

Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

The paper introduces MOSAIC, a method for personalized federated vision‑language models that addresses domain heterogeneity by focusing on class‑specific cross‑domain residuals. It constructs a decision‑aware harmfulness score to identify residuals that hurt image‑text decision margins, then uses a low‑rank residual adapter with shared class factors and private domain factors, an image‑conditioned gate, and harmful‑pair‑aware reweighting to refine local updates. Experiments on Office31, OfficeHome, and DomainNet100 show consistent improvements in macro‑client top‑1 accuracy across various domain‑shift scenarios.

By Wentao Yue, Qingyu Mao, Tianyou Lai, Ahmed M. Abdelmoniem, Qilei Li
arXiv Computation and Language
5d ago

The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

The paper introduces Percept-V, a dataset of 6,000 program-generated images across 30 domains designed to test simple visual perception skills from the TVPS-4 framework. Experiments show that state‑of‑the‑art multimodal large language models perform poorly compared to humans, especially as image complexity increases, and that fine‑tuning yields only limited generalization to related datasets. The study highlights specific perception skills that remain challenging for current models.

By Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla