A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline reduces human annotation effort and improves transcription accuracy across subsequent pages. The authors apply this method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the pipeline against leading multimodal large language models.
The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline improves transcription accuracy across subsequent pages, reducing the need for costly human annotation. The authors apply the method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the results against leading multimodal large language models.
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
arXiv:2607. 20385v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries.