HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
We present HunyuanOCR-1. 5, a lightweight end-to-end OCR-specialized vision-language model.
We present HunyuanOCR-1. 5, a lightweight end-to-end OCR-specialized vision-language model.
HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
arXiv:2605. 16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography.
arXiv:2609.36136v1 Announce Type: new Abstract: Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on vi...
Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.
arXiv:2607. 13639v1 Announce Type: cross Abstract: We introduce OvisOCR2, a 0.
The paper introduces an all‑in‑one multilingual scene text recognizer called ScriptMoE, which uses a script‑aware mixture‑of‑experts architecture to handle 10 scripts and 229 languages. It is built on a new large‑scale synthetic dataset, TextMuSS‑10M, and evaluated on the TextMuSS‑Bench, achieving 82.06% accuracy—1.31% higher than the best baseline. When integrated into the PP‑OCRv5 pipeline, ScriptMoE raises the end‑to‑end multilingual F1 score from 65.71% to 80.89%, slightly surpassing the best vision‑language model while using far fewer parameters.
The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.
This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.
arXiv:2608.30678v1 Announce Type: new Abstract: Text-rich image understanding requires multimodal large language models (MLLMs) to organize OCR (Optical Character Recognition)-grounded evidence acros...
arXiv:2609.00232v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmar...