From Pixels to Prompts: Vision-Language Models
arXiv:2605. 07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago.
Related stories
Vision Language Models Explained
Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
A Dive into Vision-Language Models
Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes
arXiv:2602. 00593v4 Announce Type: replace-cross Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy overlooked by existing benchmarks that evaluate these abilities in isolation.
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.
Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge
arXiv:2606. 27527v1 Announce Type: cross Abstract: Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models in other modalities remains underexplored.
Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
The paper introduces the Separable Law, a framework that predicts how vision‑language model performance varies with language backbone size and visual token count. By fitting this law to 26 InternVL and QwenVL models across a range of backbone sizes (1B–72B) and image resolutions (224–8K pixels), the authors show that some question types scale predictably with model capacity while others do not. The law also provides a closed‑form rule for allocating compute between backbone size and visual tokens, helping to choose near‑optimal model and image sizes under a fixed budget.
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models
arXiv:2606. 07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.
Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures
arXiv:2609.00663v1 Announce Type: new Abstract: Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separatin...
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
arXiv:2606.19100v4 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open...
The Geometry of Representational Failures in Vision Language Models
arXiv:2602. 07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing to identify the most similar objects among distractions.