arXiv AI By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

Read the original on arXiv AI →

arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

arXiv:2606. 01393v1 Announce Type: cross Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems.

By Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo, Zhenting Qi, Konwoo Kim, Longtian Ye, Xiaolong Luo, Jinhe Bi, Henry Zhang, Haris Riaz, Xuan Zhang, Yunze Xiao, Bangya Liu, Tom Tang, Yunfei Zhao, Qunshu Lin, Zihan Wang, Minghao Liu, Michael Lingzhi Li, Yilun Du, Jesse Thomason, Rogerio Feris, Alex Pentland, Zexue He
arXiv Computation and Language
4d ago

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings

The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.

By Zahra Anvari
arXiv Computer Vision
Sep 18

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.

By Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, Yu Zhou
arXiv AI
Aug 20

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.

By Elias Schubert, Felix Bie{\ss}mann
arXiv Computation and Language
Aug 24

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...

By A. Said Gurbuz (IBM Research Zurich, ETH Zurich), Ahmed Nassar (IBM Research Zurich), Christoph Auer (IBM Research Zurich), Maksym Lysak (IBM Research Zurich), Lucas Morin (IBM Research Zurich), Matteo Omenetti (IBM Research Zurich), Tim Strohmeyer (IBM Research Zurich), Panagiotis Vagenas (IBM Research Zurich), Nikolaos Livathinos (IBM Research Zurich), Michele Dolfi (IBM Research Zurich), Peter Staar (IBM Research Zurich)