arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
By Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv:2606. 07558v1 Announce Type: cross Abstract: Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting impractical at scale.
By Kateryna Lutsai, Pavel Stra\v{n}\'ak, David Nov\'ak, Dana K\v{r}iv\'ankov\'a
HERBIOME is a modular, end‑to‑end pipeline that automates the digitization of herbarium labels. It combines YOLOv8 for component detection, CRAFT Hezar for word‑level text localization, a fine‑tuned TrOCR model for mixed handwritten and printed text recognition, and GPT‑4o Mini for structuring metadata into standardized fields. Evaluation on 450 French specimens shows high surface similarity (MWS ≈ 0.616) and moderate semantic accuracy (SMA ≈ 0.442), with taxonomic fields identified as the main challenge.
By Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare, Nicolas Turenne, Eric Chenin, Abderrazak Sebaa, Edi Prifti, Jean-Daniel Zucker, Youcef Sklab
PyPottery is an open‑source, AI‑powered suite that semi‑automates the entire ceramic documentation pipeline, comprising four modules: PyPotteryScan for image extraction and handwriting recognition, PyPotteryInk for automatic inking of pencil drawings, PyPotteryTrace for semantically‑aware vectorization, and PyPotteryLayout for automated layout generation. In a study of 50 hand‑drawn sheets with 240 pottery drawings from the Terramara di Montale in Italy, users reported a median perceived speedup of 40× compared to traditional workflows, with a range from 17.5× to 120×. The results demonstrate significant time savings and suggest that AI can shift cognitive labor toward augmentation rather than full automation.
By Lorenzo Cardarelli
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline reduces human annotation effort and improves transcription accuracy across subsequent pages. The authors apply this method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the pipeline against leading multimodal large language models.
By Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline improves transcription accuracy across subsequent pages, reducing the need for costly human annotation. The authors apply the method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the results against leading multimodal large language models.
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.
By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla
arXiv:2608.23263v1 Announce Type: new
Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to pro...
By Zeyd Boukhers, Lingxiao Kong, Xenophon Zabulis, Georgios Toubekis
The paper introduces SCAM, a line-level dataset of digitized Sahidic Coptic ancient manuscripts designed for Handwritten Text Recognition in low-resource settings. SCAM captures realistic challenges such as varied acquisition conditions, ink fading, bleed-through, and material deterioration, while also presenting linguistic difficulties due to the scarce resources, uncommon alphabet, and dialect-specific diacritics of Sahidic Coptic. The authors benchmark several state‑of‑the‑art HTR methods, demonstrating the performance gap between modern, well‑resourced scripts and historically grounded, low‑resource scenarios.
By Fabio Quattrini, Carmine Zaccagnino, Costanza Bianchi, Silvia Cascianelli, Rita Cucchiara
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which ev...
arXiv:2607. 24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives.
By Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)