arXiv Computer Vision

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

arXiv AI
Jul 21

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.

By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv Machine Learning
Aug 13

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

arXiv:2605. 16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography.

By Qinwu Xu, Yifan Jiang, Haoyu Ren
arXiv Computation and Language
Aug 31

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.

By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
arXiv AI
Aug 20

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.

By Elias Schubert, Felix Bie{\ss}mann
arXiv Computer Vision
Sep 11

TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

TeleOCR is a unified framework for document parsing that tackles challenges in both decoupled and end-to-end Vision‑Language Models. It introduces deformation‑aware learning to handle geometric distortions, an adaptive sampling mechanism for complex layouts, and a content‑structure decoupled strategy to model formula grammars and table structures. The approach achieves state‑of‑the‑art results on multiple benchmarks, including top placement in the ICDAR 2026 Sci‑ImageMiner Challenge.

By Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, MingKun Jiang, Zhongjiang He, Hao Sun
arXiv Computation and Language
Aug 24

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...

By A. Said Gurbuz (IBM Research Zurich, ETH Zurich), Ahmed Nassar (IBM Research Zurich), Christoph Auer (IBM Research Zurich), Maksym Lysak (IBM Research Zurich), Lucas Morin (IBM Research Zurich), Matteo Omenetti (IBM Research Zurich), Tim Strohmeyer (IBM Research Zurich), Panagiotis Vagenas (IBM Research Zurich), Nikolaos Livathinos (IBM Research Zurich), Michele Dolfi (IBM Research Zurich), Peter Staar (IBM Research Zurich)
arXiv Machine Learning
Aug 28

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.

By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah