arXiv:2610.01134v1 Announce Type: new
Abstract: An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are a...
By Faias Satter, Sk. Md. Masudul Ahsan
arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
By Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide
The paper evaluates whether incorporating the hierarchical structure of Tironian notes can improve automatic recognition of this complex Latin shorthand system. Experiments compare flat classifiers (ResNet18, ConvNeXt, Swin, ViT) with hierarchy‑aware models (HD‑CNN and routing approaches) on handwritten and manuscript samples, with and without few‑shot adaptation. Results show that hierarchical models outperform flat ones when no adaptation is applied, but flat models surpass them after few‑shot adaptation, indicating that hierarchy can aid recognition under non‑adapted conditions.
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data.
arXiv:2609.37755v1 Announce Type: new
Abstract: Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would...
By Anton Repushko, Elena Chepel
arXiv:2604. 00725v2 Announce Type: replace-cross Abstract: End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts.
By Merveilles Agbeti-Messan, Pierrick Tranouez, St\'ephane Nicolas, Cl\'ement Chatelain, Thierry Paquet