arXiv Machine Learning By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler

Learning Diachronic Representations of Ancient Greek Letterforms

Read the original on arXiv Machine Learning →

arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
5d ago

Evaluating Hierarchy-Aware Deep Learning for the Recognition of Tironian Notes

The paper evaluates whether incorporating the hierarchical structure of Tironian notes can improve automatic recognition of this complex Latin shorthand system. Experiments compare flat classifiers (ResNet18, ConvNeXt, Swin, ViT) with hierarchy‑aware models (HD‑CNN and routing approaches) on handwritten and manuscript samples, with and without few‑shot adaptation. Results show that hierarchical models outperform flat ones when no adaptation is applied, but flat models surpass them after few‑shot adaptation, indicating that hierarchy can aid recognition under non‑adapted conditions.

arXiv Machine Learning
Sep 24

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.

By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu
arXiv Computer Vision
Aug 31

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.

By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla
arXiv Machine Learning
Jul 14

Tokenization vs. Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition

arXiv:2603. 16883v2 Announce Type: replace-cross Abstract: Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability.

By Jindong Li, Dario Zanca, Vincent Christlein, Tim Hamann, Jens Barth, Peter K\"ampf, Bj\"orn Eskofier
arXiv Computer Vision
Sep 11

Rethinking Handwritten Character Recognition

The paper introduces GraphemeNet, a unified multi‑script handwritten character recognition architecture that explicitly encodes script‑geometric regularities. It uses two orthogonal binary axes: Persistent Scaffold Injection (PSI) to embed stroke‑level geometry into each encoder stage, and a choice between gated global pooling or a Stroke Topology Module for spatial relational reasoning. Across fourteen benchmarks in eight writing systems, GraphemeNet achieves state‑of‑the‑art performance with fewer parameters, demonstrating the effectiveness of structural‑prior efficiency for multi‑script HCR.

By Ranjit Raut, Aarav Subedi, Ashim Shrestha