arXiv AI

ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.

arXiv Machine Learning
Jul 28

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

arXiv:2607. 24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives.

By Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
arXiv AI
Jul 7

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

arXiv:2607. 03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals.

By Nguyen Kim Hai Bui, Md. Easin Arafat, Tam\'as G\'abor Orosz, Mufti Mahmud
arXiv Computation and Language
Sep 3

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.

By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova
arXiv AI
Aug 11

TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents

arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv AI
Aug 12

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv Computation and Language
4d ago

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings

The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.

By Zahra Anvari
arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv Computer Vision
Sep 18

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

The paper adapts the compact PP‑OCRv6 recognizer for historical text recognition and compares it to a conventional CRNN across various training regimes, including generalized pretraining, domain‑specific training, corpus‑level fine‑tuning, and manuscript‑specific few‑shot adaptation on multilingual Latin and Arabic scripts. While PP‑OCRv6 does not always beat the CRNN when trained from scratch, heterogeneous pretraining significantly improves its generalization. Additionally, fine‑tuned PP‑OCRv6 can surpass a large vision‑language model (Qwen3.5‑based Medusa) that is specifically tailored for historical Latin‑script handwriting recognition.

By Benjamin Kiessling (ALMAnaCH)