The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.
By Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li, Junlin Jiang
The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.
By Zahra Anvari
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
We present HunyuanOCR-1. 5, a lightweight end-to-end OCR-specialized vision-language model.
arXiv:2609.01575v1 Announce Type: new
Abstract: Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a...
By Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin, Olga Tsymboi, Anatolii Potapov, Aleksandr Ivanov
arXiv:2607. 03650v1 Announce Type: cross Abstract: Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
By Enshuo Hsu, Jin Zhou, Kirk Roberts
HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.
By Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, Yu Zhou
arXiv:2608.20868v1 Announce Type: cross
Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...
By A. Said Gurbuz (IBM Research Zurich, ETH Zurich), Ahmed Nassar (IBM Research Zurich), Christoph Auer (IBM Research Zurich), Maksym Lysak (IBM Research Zurich), Lucas Morin (IBM Research Zurich), Matteo Omenetti (IBM Research Zurich), Tim Strohmeyer (IBM Research Zurich), Panagiotis Vagenas (IBM Research Zurich), Nikolaos Livathinos (IBM Research Zurich), Michele Dolfi (IBM Research Zurich), Peter Staar (IBM Research Zurich)
This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.
By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah
arXiv:2609.37712v1 Announce Type: cross
Abstract: Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, loca...
By GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li, Jing Huang, Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu, Xiang Qi, Weiqiang Wang
Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.
The paper introduces a benchmark for evaluating open‑source instruction‑tuned large language models (LLMs) on key‑value pair extraction from documents under both clean‑text and noisy OCR conditions. It tests decoder‑only models such as Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek on FUNSD, CORD, and SROIE datasets, using OCR outputs from PaddleOCR, EasyOCR, and Tesseract. Results show that while modern LLMs perform well on high‑quality text, their performance drops sharply with OCR noise, and the main determinants of success are semantic reasoning and textual fidelity, with larger models offering diminishing returns under noisy inputs.
By Zahra Anvari, Vassilis Athitsos