Finetuning olmOCR to be a faithful OCR-Engine
Related stories
SOTA OCR with Core ML and dots.ocr
Introducing Mistral OCR 3
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
arXiv:2607. 03650v1 Announce Type: cross Abstract: Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs).
Mistral OCR
PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
arXiv:2609.37712v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, loca...
When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
arXiv:2608.22366v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remai...
OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement
OCR-EDR is a rendering‑aware framework that diagnoses and repairs OCR errors by jointly evaluating an OCR prediction, its editable form, and the rendered image of the source document. It localizes genuine mistakes while preserving valid or rendering‑equivalent predictions, then applies executable edits and iteratively reassesses with updated renderings. On the newly constructed OCRErrBench, the DocEDR model achieves 94.78% diagnostic accuracy and repairs 86.23% of errors, boosting formula metrics by up to 30.99 percentage points and improving CDM scores on several OCR systems.
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.
Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data.
Color Independent Word Segmentation From Transcribed Bangla Passages
The paper presents a color‑independent word segmentation method for handwritten Bangla text images. It works on smartphone‑captured images regardless of paper color or ink type, and the custom dataset includes various real‑world challenges such as shadows. The system achieves 90.60 % recall, 91.80 % precision, and 91.20 % F1‑score on 7,374 words.
UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.