arXiv Machine Learning
Sep 14

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.

By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
arXiv Computer Vision
Sep 18

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

The paper adapts the compact PP‑OCRv6 recognizer for historical text recognition and compares it to a conventional CRNN across various training regimes, including generalized pretraining, domain‑specific training, corpus‑level fine‑tuning, and manuscript‑specific few‑shot adaptation on multilingual Latin and Arabic scripts. While PP‑OCRv6 does not always beat the CRNN when trained from scratch, heterogeneous pretraining significantly improves its generalization. Additionally, fine‑tuned PP‑OCRv6 can surpass a large vision‑language model (Qwen3.5‑based Medusa) that is specifically tailored for historical Latin‑script handwriting recognition.

By Benjamin Kiessling (ALMAnaCH)
arXiv Computer Vision
Aug 31

What Can Low Resource Languages Learn From Each Other?

The paper examines OCR adaptation for low‑resource languages, noting that fine‑tuning often hits a performance ceiling in data‑scarce settings. It identifies that lower layers of language‑specific models learn redundant features while higher layers capture script nuances, leading to a structural inefficiency. To address this, the authors propose PSMC, a framework that pre‑trains a base model, specializes it per language, merges the experts via task arithmetic, and co‑trains a unified multilingual backbone, achieving about a 2% improvement in Word Recognition Rate across 10 Indian scripts without adding parameters.

By Achyuth P, Kahaan Shah, Chetan Arora