arXiv Computer Vision

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.

arXiv Machine Learning
Aug 28

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

This systematic literature review examines 97 studies on optical character recognition (OCR) from 2015 to 2025, tracing the evolution of AI models, application domains, data types, and linguistic coverage. It identifies key OCR models, evaluates their performance, strengths, and limitations, and highlights unresolved challenges such as limited resources for underrepresented languages, high variability in handwritten text, and constraints in real‑time applications. The review proposes promising approaches—including self‑supervised learning, multimodal AI, AutoML, AI‑assisted postprocessing, TinyML, and joint corpora creation—to enhance OCR accuracy and address these challenges for industrial use.

By Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht, Mohd Shahrizal Rusli, Shahidatul Sadiah
arXiv AI
Aug 11

TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents

arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv AI
Aug 12

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
Hugging Face Trending Papers
Aug 19

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline improves transcription accuracy across subsequent pages, reducing the need for costly human annotation. The authors apply the method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the results against leading multimodal large language models.

arXiv AI
Aug 20

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

The paper presents a local traditional OCR pipeline that can be iteratively fine‑tuned on both layout and appearance levels of complex historical Sanskrit manuscripts. By adapting to the specific manuscript distribution, the pipeline reduces human annotation effort and improves transcription accuracy across subsequent pages. The authors apply this method to three manuscripts, release a dataset with detailed layout and Unicode annotations in PAGE‑XML format, and benchmark the pipeline against leading multimodal large language models.

By Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
arXiv AI
Aug 20

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.

By Elias Schubert, Felix Bie{\ss}mann
arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv AI
Aug 20

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.

By Zinuo Guo, Min Zhang, Bo Jiang