arXiv AI

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.

arXiv AI
Sep 15

Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction

The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.

By Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li, Junlin Jiang
arXiv Computation and Language
6d ago

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings

The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.

By Zahra Anvari
arXiv AI
Jul 21

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.

By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv Computation and Language
Sep 2

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

arXiv:2609.01575v1 Announce Type: new Abstract: Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a...

By Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin, Olga Tsymboi, Anatolii Potapov, Aleksandr Ivanov
arXiv Computer Vision
Sep 18

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.

By Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, Yu Zhou
arXiv Computation and Language
Aug 24

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...

By A. Said Gurbuz (IBM Research Zurich, ETH Zurich), Ahmed Nassar (IBM Research Zurich), Christoph Auer (IBM Research Zurich), Maksym Lysak (IBM Research Zurich), Lucas Morin (IBM Research Zurich), Matteo Omenetti (IBM Research Zurich), Tim Strohmeyer (IBM Research Zurich), Panagiotis Vagenas (IBM Research Zurich), Nikolaos Livathinos (IBM Research Zurich), Michele Dolfi (IBM Research Zurich), Peter Staar (IBM Research Zurich)
arXiv Machine Learning
Aug 28

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.

By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah
arXiv AI
6d ago

PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

arXiv:2609.37712v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, loca...

By GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li, Jing Huang, Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu, Xiang Qi, Weiqiang Wang
Hugging Face Trending Papers
Jun 24

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.

arXiv Computation and Language
Sep 17

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

The paper introduces a benchmark for evaluating open‑source instruction‑tuned large language models (LLMs) on key‑value pair extraction from documents under both clean‑text and noisy OCR conditions. It tests decoder‑only models such as Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek on FUNSD, CORD, and SROIE datasets, using OCR outputs from PaddleOCR, EasyOCR, and Tesseract. Results show that while modern LLMs perform well on high‑quality text, their performance drops sharply with OCR noise, and the main determinants of success are semantic reasoning and textual fidelity, with larger models offering diminishing returns under noisy inputs.

By Zahra Anvari, Vassilis Athitsos