arXiv Computer Vision

Rethinking Handwritten Character Recognition

The paper introduces GraphemeNet, a unified multi‑script handwritten character recognition architecture that explicitly encodes script‑geometric regularities. It uses two orthogonal binary axes: Persistent Scaffold Injection (PSI) to embed stroke‑level geometry into each encoder stage, and a choice between gated global pooling or a Stroke Topology Module for spatial relational reasoning. Across fourteen benchmarks in eight writing systems, GraphemeNet achieves state‑of‑the‑art performance with fewer parameters, demonstrating the effectiveness of structural‑prior efficiency for multi‑script HCR.

arXiv Computer Vision
1d ago

Hybrid Feature Learning for Handwriting Verification

The paper introduces a Hybrid Deep Learning (HDL) architecture that combines Auto-Learned Features (ALF) and Human-Engineered Features (HEF) for handwriting verification. ALF is extracted using a Two Channel Convolutional Neural Network (TC-CNN) or a Two Channel Autoencoder (TC-AE), while HEF is obtained via Gradient Structural Concavity (GSC) or Scale Invariant Feature Transform (SIFT). Experiments on 150,000 pairs of the word "AND" from 1,500 writers show that the HDL model using AE-GSC achieves 99.7% accuracy on a seen writer dataset and 92.16% on a shuffled writer dataset, outperforming CEDAR-FOX, and AE-SIFT performs comparably on unseen writers.

By Seyed Mohammad Abuzar Hashemi, Mihir Chauhan, Jun Chu, Sargur Srihari
arXiv Machine Learning
Jun 25

Learning Diachronic Representations of Ancient Greek Letterforms

arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.

By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
arXiv Computer Vision
Aug 31

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.

By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla
arXiv Machine Learning
Sep 10

Rotation-free Online Handwritten Character Recognition Using Linear Recurrent Units

The paper presents a rotation‑free online handwritten character recognition system that uses Sliding Window Path Signature (SW‑PS) to extract local structural features and a lightweight Linear Recurrent Unit (LRU) classifier. The LRU blends the incremental processing of RNNs with the parallel training efficiency of state‑space models to model dynamic stroke characteristics. Experiments on rotated CASIA‑OLHWDB1.1 subsets (digits, English upper letters, Chinese radicals) achieved accuracies of 99.62%, 96.67%, and 94.33% respectively, outperforming competing models in convergence speed and test accuracy.

By Zhe Ling, Sicheng Yu, Danyu Yang
arXiv Machine Learning
Jul 14

Tokenization vs. Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition

arXiv:2603. 16883v2 Announce Type: replace-cross Abstract: Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability.

By Jindong Li, Dario Zanca, Vincent Christlein, Tim Hamann, Jens Barth, Peter K\"ampf, Bj\"orn Eskofier
arXiv AI
Aug 20

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.

By Zinuo Guo, Min Zhang, Bo Jiang
arXiv Machine Learning
Sep 14

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.

By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
arXiv Computer Vision
Sep 22

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

The paper introduces an all‑in‑one multilingual scene text recognizer called ScriptMoE, which uses a script‑aware mixture‑of‑experts architecture to handle 10 scripts and 229 languages. It is built on a new large‑scale synthetic dataset, TextMuSS‑10M, and evaluated on the TextMuSS‑Bench, achieving 82.06% accuracy—1.31% higher than the best baseline. When integrated into the PP‑OCRv5 pipeline, ScriptMoE raises the end‑to‑end multilingual F1 score from 65.71% to 80.89%, slightly surpassing the best vision‑language model while using far fewer parameters.

By Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen