Introducing Mistral OCR 3
Related stories
Introducing Mistral OCR 4
Mistral OCR 4 delivers enterprise document AI with 170-language support, bounding boxes, and self-hosted deployment.
Finetuning olmOCR to be a faithful OCR-Engine
SOTA OCR with Core ML and dots.ocr
Introducing Mistral 3
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
arXiv:2607. 03650v1 Announce Type: cross Abstract: Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
Supercharge your OCR Pipelines with Open Models
PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters
Introducing Mistral Code
PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.
Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data.