OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.
By Zinuo Guo, Min Zhang, Bo Jiang
arXiv:2606. 08858v1 Announce Type: cross Abstract: The automatic processing of handwritten forms remains a challenging task, wherein detection and subsequent classification of handwritten characters are essential steps.
By Hartwig Grabowski
arXiv:2603. 16883v2 Announce Type: replace-cross Abstract: Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability.
By Jindong Li, Dario Zanca, Vincent Christlein, Tim Hamann, Jens Barth, Peter K\"ampf, Bj\"orn Eskofier
Teaching machines to emulate natural handwriting styles remains an open challenge, as it requires synthesizing stroke sequences that dynamically vary in shape, texture, pressure and script - not only across individuals, but also within a single person's handwriting. Attempts at this challenge have largely explored deep learning methods in both online and offline settings.
The paper introduces GraphemeNet, a unified multi‑script handwritten character recognition architecture that explicitly encodes script‑geometric regularities. It uses two orthogonal binary axes: Persistent Scaffold Injection (PSI) to embed stroke‑level geometry into each encoder stage, and a choice between gated global pooling or a Stroke Topology Module for spatial relational reasoning. Across fourteen benchmarks in eight writing systems, GraphemeNet achieves state‑of‑the‑art performance with fewer parameters, demonstrating the effectiveness of structural‑prior efficiency for multi‑script HCR.
By Ranjit Raut, Aarav Subedi, Ashim Shrestha
The paper introduces PA-CDM, a new position‑aware character detection matching metric for handwritten mathematical expression recognition that improves on existing render‑based and tree‑edit metrics by incorporating position‑forest encoding and divergence‑level weighting. It also presents StructPerturb v2.0, a benchmark of 1,340 controlled perturbation pairs, and a cross‑metric consistency protocol that includes a sensitivity matrix, a human study, and LLM‑judge calibration. In a six‑annotator study, PA‑CDM achieves the highest correlation with human judgments (Spearman rho = 0.9535) among seven automatic metrics, approaching the performance of a costly LLM judge while remaining deterministic and cost‑free.
By Shiliang Luo (East China Normal University)
arXiv:2607. 15509v1 Announce Type: cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition.
By Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
arXiv:2609.05662v1 Announce Type: cross
Abstract: Full-page end-to-end Optical Music Recognition seeks to transcribe entire music pages directly into symbolic notation, avoiding the limitations of tr...
By Adrian Rosello, Antonio R\'ios-Vila, David Rizo, Jorge Calvo-Zaragoza
arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.
By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
The paper introduces a Hybrid Deep Learning (HDL) architecture that combines Auto-Learned Features (ALF) and Human-Engineered Features (HEF) for handwriting verification. ALF is extracted using a Two Channel Convolutional Neural Network (TC-CNN) or a Two Channel Autoencoder (TC-AE), while HEF is obtained via Gradient Structural Concavity (GSC) or Scale Invariant Feature Transform (SIFT). Experiments on 150,000 pairs of the word "AND" from 1,500 writers show that the HDL model using AE-GSC achieves 99.7% accuracy on a seen writer dataset and 92.16% on a shuffled writer dataset, outperforming CEDAR-FOX, and AE-SIFT performs comparably on unseen writers.
By Seyed Mohammad Abuzar Hashemi, Mihir Chauhan, Jun Chu, Sargur Srihari
arXiv:2609.37195v1 Announce Type: new
Abstract: Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down ti...
By Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza
WildHandBench is a new benchmark comprising 500 handwritten documents that span free text, tables, and formulas across four languages and nine real‑world scenarios. It introduces a Prior‑Driven Error (PDE) metric to assess whether mistakes stem from language priors rather than visual cues. In tests of 18 state‑of‑the‑art models, the best achieves only 71.85% accuracy, while humans reach 77.09%, and model errors are largely prior‑driven (63‑91%) compared to human errors (49%).
By Jun Zhang, Qiao Zhao, Cheng Cui, Jianying Qu, Zhongkai Sun, Jianwen Yang, Changda Zhou, ZhuoXin Liu, Shubin Han