Unlocking the potential of vision language models on satellite imagery through fine-tuning
Related stories
Vision Language Models (Better, faster, stronger)
Accelerating Vision-Language Models: BridgeTower on Habana Gaudi2
Vision Language Models Explained
A Dive into Vision-Language Models
PaliGemma – Google's Cutting-Edge Open Vision Language Model
Accelerating vision-language models with LFM2.5-VL-DSpark
SmolVLM - small yet mighty Vision Language Model
Welcome PaliGemma 2 – New vision language models by Google
SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding
arXiv:2605.17949v2 Announce Type: replace Abstract: Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the...
Preference Optimization for Vision Language Models
Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
The paper "Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers" reports that vision‑language models (VLMs) struggle to read images containing two overlapping text layers—one with sharp contour lines and one with soft shading. Using the DecoyBench dataset of 300 such images, the authors evaluated six closed‑source VLMs under naive and guided prompting at high and low resolutions. While humans could read both layers accurately, the models reliably read only the contour layer at high resolution and failed to extract the shading layer; at low resolution, neither the models nor humans could read the contour layer, but the shading layer remained readable. "whyItMatters":"The study highlights a consistent limitation of current VLMs in handling typographic structures with multiple spatial frequency layers, underscoring their vulnerability to typographic attacks and the need for more robust text‑recognition capabilities."