Mistral AI

Unlocking the potential of vision language models on satellite imagery through fine-tuning

arXiv AI
Sep 28

Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers

The paper "Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers" reports that vision‑language models (VLMs) struggle to read images containing two overlapping text layers—one with sharp contour lines and one with soft shading. Using the DecoyBench dataset of 300 such images, the authors evaluated six closed‑source VLMs under naive and guided prompting at high and low resolutions. While humans could read both layers accurately, the models reliably read only the contour layer at high resolution and failed to extract the shading layer; at low resolution, neither the models nor humans could read the contour layer, but the shading layer remained readable. "whyItMatters":"The study highlights a consistent limitation of current VLMs in handling typographic structures with multiple spatial frequency layers, underscoring their vulnerability to typographic attacks and the need for more robust text‑recognition capabilities."

By Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an