Apollo Restore is a 24‑billion‑parameter large language model fine‑tuned from Mistral Small to fill in gaps in fragmentary Ancient Greek texts using a fill‑in‑the‑middle objective. It achieves state‑of‑the‑art performance on short lacunae, placing the correct restoration among its top twenty candidates for 80.6% of documentary‑papyrus, 54.6% of literary‑papyrus, and 61.0% of stone‑inscription gaps, outperforming previous models by significant margins. In blind expert evaluations, 20 specialists preferred Apollo Restore over the strongest baseline and judged its performance at least as good as human restorations in 77% of cases, while also improving the published reading of the Vesuvius‑carbonised papyrus P.Herc. 1667.
By Hope McGovern, Anna Dolganov, Samuel Belkadi, Guillaume Kunsch, Dimitris Vlitas, David A. Smith
arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
By Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide
arXiv:2609.37755v1 Announce Type: new
Abstract: Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would...
By Anton Repushko, Elena Chepel
The paper introduces SCAM, a line-level dataset of digitized Sahidic Coptic ancient manuscripts designed for Handwritten Text Recognition in low-resource settings. SCAM captures realistic challenges such as varied acquisition conditions, ink fading, bleed-through, and material deterioration, while also presenting linguistic difficulties due to the scarce resources, uncommon alphabet, and dialect-specific diacritics of Sahidic Coptic. The authors benchmark several state‑of‑the‑art HTR methods, demonstrating the performance gap between modern, well‑resourced scripts and historically grounded, low‑resource scenarios.
By Fabio Quattrini, Carmine Zaccagnino, Costanza Bianchi, Silvia Cascianelli, Rita Cucchiara
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
The paper introduces a retrieval-guided glyph-aware restoration framework for low-resource Manchu historical documents. Unlike traditional pixel-level reconstruction methods, it incorporates glyph-level structural knowledge by retrieving relevant glyph exemplars to guide the restoration process. Experiments show that this approach improves both image quality and glyph fidelity compared to existing methods.
By Ting Huang, Dongdong Wang, Mingqiu Liang, Siyang Lu