arXiv AI

Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers

The paper "Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers" reports that vision‑language models (VLMs) struggle to read images containing two overlapping text layers—one with sharp contour lines and one with soft shading. Using the DecoyBench dataset of 300 such images, the authors evaluated six closed‑source VLMs under naive and guided prompting at high and low resolutions. While humans could read both layers accurately, the models reliably read only the contour layer at high resolution and failed to extract the shading layer; at low resolution, neither the models nor humans could read the contour layer, but the shading layer remained readable. "whyItMatters":"The study highlights a consistent limitation of current VLMs in handling typographic structures with multiple spatial frequency layers, underscoring their vulnerability to typographic attacks and the need for more robust text‑recognition capabilities."

arXiv Machine Learning
Aug 28

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.

By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah
arXiv Computer Vision
Aug 24

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt

The paper introduces Embedded Stroop, a diagnostic test that embeds text prompts directly into images to study interference in multimodal large language models (MLLMs). Using the What-Color-Is-the-Text (WCIT) benchmark, which tests 59 fine‑grained colors in standard, flipped, and masked conditions, the authors evaluate 16 models and find that while exact color accuracy is low (6.3%), models still recognize coarse color families (38.4%) but frequently hallucinate the embedded word instead of the true color (Stroop Hallucination Rate of 21.6%). Masking or flipping the embedded text reduces hallucinations, indicating that semantic legibility can dominate visual color perception in MLLMs.

By Jinkun Zhao, Lei Huang, Haixin Ge, Wenjun Wu
Hugging Face Trending Papers
Jun 24

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.

arXiv AI
Sep 17

Using OCR Heads to Verbalize Image Semantics

The paper investigates how vision‑language models (VLMs) perform optical character recognition (OCR) by identifying attention heads that are causally necessary for OCR across four models. These heads are shown to be general‑purpose, producing interpretable semantic features for any image token, such as recognizing the word "bike" or the concept "feathers". By collapsing the heads’ attention weights into a verbalization lens transformation, the authors reveal that image representations align with language from early layers and can even be used to edit non‑word concepts in images, demonstrating the broader utility of this subspace.

By Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau
arXiv Computation and Language
Sep 25

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.

By Cong Xu, Ravi Sankar
arXiv AI
Sep 3

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

Vision‑Language Models (VLMs) are increasingly replacing traditional OCR for document understanding, but this study shows they often rewrite imperfect text into more plausible forms, a flaw that clean‑text OCR benchmarks miss. The authors created FaithC4, a multilingual perturbation benchmark of 1,455 single‑page documents with scramble, random substitution, and visually similar substitution attacks, and evaluated 15 systems across general‑purpose VLMs, OCR‑specialized VLMs, and traditional OCR pipelines. Results reveal that general‑purpose VLMs suffer up to 6.9 WER points under perturbation, OCR‑specialized VLMs 0.1–3.4 points, and traditional OCR less than 0.8 points on English; probing Qwen3‑VL‑4B shows rewriting occurs only when a perturbed word’s final‑layer representation remains close to the original, with short words (4–6 characters) rewritten up to 10% of the time. whyItMatters":"The findings highlight a critical limitation of VLMs in document transcription, underscoring the need for robust evaluation benchmarks that capture rewriting behavior beyond clean‑text accuracy."

By Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis