arXiv:2609.13815v1 Announce Type: new
Abstract: Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a task made substantially harder when images...
By Ritali Vatsi, Rachapudi Jagadeesh, Shruti Singh Baghel, Himani Sharma, Amit Shukla, Pawan Goyal
arXiv:2605.27750v2 Announce Type: replace-cross
Abstract: Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually uns...
By Antonia Karamolegkou, Nicolas Angleraud, Beno\^it Sagot, Thibault Cl\'erice
arXiv:2609.00232v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmar...
By Yue Zhou, Yuan Wu, Yi Chang
This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.
By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah
The paper introduces Embedded Stroop, a diagnostic test that embeds text prompts directly into images to study interference in multimodal large language models (MLLMs). Using the What-Color-Is-the-Text (WCIT) benchmark, which tests 59 fine‑grained colors in standard, flipped, and masked conditions, the authors evaluate 16 models and find that while exact color accuracy is low (6.3%), models still recognize coarse color families (38.4%) but frequently hallucinate the embedded word instead of the true color (Stroop Hallucination Rate of 21.6%). Masking or flipping the embedded text reduces hallucinations, indicating that semantic legibility can dominate visual color perception in MLLMs.
By Jinkun Zhao, Lei Huang, Haixin Ge, Wenjun Wu
Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.
arXiv:2610.01134v1 Announce Type: new
Abstract: An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are a...
By Faias Satter, Sk. Md. Masudul Ahsan
The paper investigates how vision‑language models (VLMs) perform optical character recognition (OCR) by identifying attention heads that are causally necessary for OCR across four models. These heads are shown to be general‑purpose, producing interpretable semantic features for any image token, such as recognizing the word "bike" or the concept "feathers". By collapsing the heads’ attention weights into a verbalization lens transformation, the authors reveal that image representations align with language from early layers and can even be used to edit non‑word concepts in images, demonstrating the broader utility of this subspace.
By Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau
The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.
By Cong Xu, Ravi Sankar
Vision‑Language Models (VLMs) are increasingly replacing traditional OCR for document understanding, but this study shows they often rewrite imperfect text into more plausible forms, a flaw that clean‑text OCR benchmarks miss. The authors created FaithC4, a multilingual perturbation benchmark of 1,455 single‑page documents with scramble, random substitution, and visually similar substitution attacks, and evaluated 15 systems across general‑purpose VLMs, OCR‑specialized VLMs, and traditional OCR pipelines. Results reveal that general‑purpose VLMs suffer up to 6.9 WER points under perturbation, OCR‑specialized VLMs 0.1–3.4 points, and traditional OCR less than 0.8 points on English; probing Qwen3‑VL‑4B shows rewriting occurs only when a perturbed word’s final‑layer representation remains close to the original, with short words (4–6 characters) rewritten up to 10% of the time.
whyItMatters":"The findings highlight a critical limitation of VLMs in document transcription, underscoring the need for robust evaluation benchmarks that capture rewriting behavior beyond clean‑text accuracy."
By Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis
arXiv:2610.02880v1 Announce Type: new
Abstract: Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this...
By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu