arXiv AI

From Pixels to Prompts: Vision-Language Models

arXiv:2605. 07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago.

arXiv Machine Learning
Jun 15

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

arXiv:2602. 00593v4 Announce Type: replace-cross Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation.

By Yifan Jiang, Cong Zhang, Bofei Zhang, Qiaofeng Zheng, Yifan Yang, Bingzhang Wang, Yew-Soon Ong
arXiv Computer Vision
Sep 4

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.

By Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
arXiv AI
4d ago

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

The paper introduces the Separable Law, a framework that predicts how vision‑language model performance varies with language backbone size and visual token count. By fitting this law to 26 InternVL and QwenVL models across a range of backbone sizes (1B–72B) and image resolutions (224–8K pixels), the authors show that some question types scale predictably with model capacity while others do not. The law also provides a closed‑form rule for allocating compute between backbone size and visual tokens, helping to choose near‑optimal model and image sizes under a fixed budget.

By Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li
arXiv Computer Vision
Sep 15

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

arXiv:2606.19100v4 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open...

By Diogo Gl\'oria-Silva, Jo\~ao Cardeira, Manuel Letras da Luz, Afonso Simpl\'icio, Gon\c{c}alo Vinagre, Diogo Tavares, Rafael Ferreira, In\^es Calvo, In\^es Vieira, David Semedo, Jo\~ao Magalh\~aes