arXiv Machine Learning By Merveilles Agbeti-Messan, Pierrick Tranouez, St\'ephane Nicolas, Cl\'ement Chatelain, Thierry Paquet

A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR

Read the original on arXiv Machine Learning →

arXiv:2604. 00725v2 Announce Type: replace-cross Abstract: End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

The paper investigates how to combine synthetic and real historical images to improve OCR for the endangered Manchu language. Using 60,000 synthetic and 20,306 real word images, the authors evaluate three vision‑language models and a compact CRNN across synthetic‑only, real‑only, joint, and sequential training regimes. Adding real data boosts word accuracy to 95–96%, and ensembling the best recognizers raises it to 98.27% without further training.

By Yan Hon Michael Chung, Hanlin Wang
Hugging Face Trending Papers
Sep 10

Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

The paper investigates how to best combine synthetic and real historical Manchu word images for low‑resource OCR. Using 60,000 synthetic and 20,306 real images, the authors compare three pretrained vision‑language models and a compact CRNN across synthetic‑only, real‑only, joint, and sequential training regimes. Adding real data boosts word accuracy to 95–96%, and ensembling the best recognizers raises it to 98.27% without extra training.

arXiv Computer Vision
Sep 18

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

The paper adapts the compact PP‑OCRv6 recognizer for historical text recognition and compares it to a conventional CRNN across various training regimes, including generalized pretraining, domain‑specific training, corpus‑level fine‑tuning, and manuscript‑specific few‑shot adaptation on multilingual Latin and Arabic scripts. While PP‑OCRv6 does not always beat the CRNN when trained from scratch, heterogeneous pretraining significantly improves its generalization. Additionally, fine‑tuned PP‑OCRv6 can surpass a large vision‑language model (Qwen3.5‑based Medusa) that is specifically tailored for historical Latin‑script handwriting recognition.

By Benjamin Kiessling (ALMAnaCH)
arXiv AI
Aug 12

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv AI
Aug 11

TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents

arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin