arXiv Computation and Language By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza

Text Restoration of Ancient Documents with Language Models

Read the original on arXiv Computation and Language →

The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 22

Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts

Apollo Restore is a 24‑billion‑parameter large language model fine‑tuned from Mistral Small to fill in gaps in fragmentary Ancient Greek texts using a fill‑in‑the‑middle objective. It achieves state‑of‑the‑art performance on short lacunae, placing the correct restoration among its top twenty candidates for 80.6% of documentary‑papyrus, 54.6% of literary‑papyrus, and 61.0% of stone‑inscription gaps, outperforming previous models by significant margins. In blind expert evaluations, 20 specialists preferred Apollo Restore over the strongest baseline and judged its performance at least as good as human restorations in 77% of cases, while also improving the published reading of the Vesuvius‑carbonised papyrus P.Herc. 1667.

By Hope McGovern, Anna Dolganov, Samuel Belkadi, Guillaume Kunsch, Dimitris Vlitas, David A. Smith
arXiv Computer Vision
Sep 2

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

The paper introduces SCAM, a line-level dataset of digitized Sahidic Coptic ancient manuscripts designed for Handwritten Text Recognition in low-resource settings. SCAM captures realistic challenges such as varied acquisition conditions, ink fading, bleed-through, and material deterioration, while also presenting linguistic difficulties due to the scarce resources, uncommon alphabet, and dialect-specific diacritics of Sahidic Coptic. The authors benchmark several state‑of‑the‑art HTR methods, demonstrating the performance gap between modern, well‑resourced scripts and historically grounded, low‑resource scenarios.

By Fabio Quattrini, Carmine Zaccagnino, Costanza Bianchi, Silvia Cascianelli, Rita Cucchiara
arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv AI
5d ago

Beyond Pixel Reconstruction: Retrieval-Guided Glyph-Aware Restoration for Low-Resource Manchu Historical Documents

The paper introduces a retrieval-guided glyph-aware restoration framework for low-resource Manchu historical documents. Unlike traditional pixel-level reconstruction methods, it incorporates glyph-level structural knowledge by retrieving relevant glyph exemplars to guide the restoration process. Experiments show that this approach improves both image quality and glyph fidelity compared to existing methods.

By Ting Huang, Dongdong Wang, Mingqiu Liang, Siyang Lu