arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.
By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
arXiv:2606. 07558v1 Announce Type: cross Abstract: Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting impractical at scale.
By Kateryna Lutsai, Pavel Stra\v{n}\'ak, David Nov\'ak, Dana K\v{r}iv\'ankov\'a
The paper introduces a Hybrid Deep Learning (HDL) architecture that combines Auto-Learned Features (ALF) and Human-Engineered Features (HEF) for handwriting verification. ALF is extracted using a Two Channel Convolutional Neural Network (TC-CNN) or a Two Channel Autoencoder (TC-AE), while HEF is obtained via Gradient Structural Concavity (GSC) or Scale Invariant Feature Transform (SIFT). Experiments on 150,000 pairs of the word "AND" from 1,500 writers show that the HDL model using AE-GSC achieves 99.7% accuracy on a seen writer dataset and 92.16% on a shuffled writer dataset, outperforming CEDAR-FOX, and AE-SIFT performs comparably on unseen writers.
By Seyed Mohammad Abuzar Hashemi, Mihir Chauhan, Jun Chu, Sargur Srihari
ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.
By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
arXiv:2604. 00725v2 Announce Type: replace-cross Abstract: End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts.
By Merveilles Agbeti-Messan, Pierrick Tranouez, St\'ephane Nicolas, Cl\'ement Chatelain, Thierry Paquet
arXiv:2609.10084v1 Announce Type: cross
Abstract: Generalized zero-shot learning (GZSL) has emerged as an important paradigm for visual recognition systems that must generalize to classes that were n...
By Clarence Chew, Gim Siang Chia, Sukalpa Chanda, Subhroshekhar Ghosh, Soumendu Sundar Mukherjee