Supercharge your OCR Pipelines with Open Models
Related stories
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
arXiv:2607. 03650v1 Announce Type: cross Abstract: Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
SOTA OCR with Core ML and dots.ocr
PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend
Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics
arXiv:2609.01575v1 Announce Type: new Abstract: Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a...
The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.
Developing an OCR model for Extracting Information from Invoices with Korean Language
The paper presents an OCR model designed to extract key information from Korean-language invoices. It combines deep learning with image preprocessing techniques and achieves an 87% F1‑score on a diverse set of collected invoices while maintaining negligible processing time.
UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
We present HunyuanOCR-1. 5, a lightweight end-to-end OCR-specialized vision-language model.
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.