arXiv Machine Learning By Takaya Kawakatsu, Ryo Ishiyama

GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition

Read the original on arXiv Machine Learning →

arXiv:2602. 03370v2 Announce Type: replace-cross Abstract: Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle with exposure bias and syntax inconsistency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 20

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.

By Zinuo Guo, Min Zhang, Bo Jiang
arXiv Machine Learning
Jul 14

Tokenization vs. Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition

arXiv:2603. 16883v2 Announce Type: replace-cross Abstract: Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability.

By Jindong Li, Dario Zanca, Vincent Christlein, Tim Hamann, Jens Barth, Peter K\"ampf, Bj\"orn Eskofier
Hugging Face Trending Papers
Jun 17

HandwritingAgent: Language-Driven Handwriting Synthesis in Scalable Vector Space

Teaching machines to emulate natural handwriting styles remains an open challenge, as it requires synthesizing stroke sequences that dynamically vary in shape, texture, pressure and script - not only across individuals, but also within a single person's handwriting. Attempts at this challenge have largely explored deep learning methods in both online and offline settings.

arXiv Computer Vision
Sep 11

Rethinking Handwritten Character Recognition

The paper introduces GraphemeNet, a unified multi‑script handwritten character recognition architecture that explicitly encodes script‑geometric regularities. It uses two orthogonal binary axes: Persistent Scaffold Injection (PSI) to embed stroke‑level geometry into each encoder stage, and a choice between gated global pooling or a Stroke Topology Module for spatial relational reasoning. Across fourteen benchmarks in eight writing systems, GraphemeNet achieves state‑of‑the‑art performance with fewer parameters, demonstrating the effectiveness of structural‑prior efficiency for multi‑script HCR.

By Ranjit Raut, Aarav Subedi, Ashim Shrestha
arXiv Machine Learning
Sep 14

PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition

The paper introduces PA-CDM, a new position‑aware character detection matching metric for handwritten mathematical expression recognition that improves on existing render‑based and tree‑edit metrics by incorporating position‑forest encoding and divergence‑level weighting. It also presents StructPerturb v2.0, a benchmark of 1,340 controlled perturbation pairs, and a cross‑metric consistency protocol that includes a sensitivity matrix, a human study, and LLM‑judge calibration. In a six‑annotator study, PA‑CDM achieves the highest correlation with human judgments (Spearman rho = 0.9535) among seven automatic metrics, approaching the performance of a costly LLM judge while remaining deterministic and cost‑free.

By Shiliang Luo (East China Normal University)