arXiv Computer Vision

Exploring In-Context Learning for Handwritten Text Recognition

arXiv Machine Learning
Sep 14

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.

By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
arXiv Computer Vision
Sep 18

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

The paper adapts the compact PP‑OCRv6 recognizer for historical text recognition and compares it to a conventional CRNN across various training regimes, including generalized pretraining, domain‑specific training, corpus‑level fine‑tuning, and manuscript‑specific few‑shot adaptation on multilingual Latin and Arabic scripts. While PP‑OCRv6 does not always beat the CRNN when trained from scratch, heterogeneous pretraining significantly improves its generalization. Additionally, fine‑tuned PP‑OCRv6 can surpass a large vision‑language model (Qwen3.5‑based Medusa) that is specifically tailored for historical Latin‑script handwriting recognition.

By Benjamin Kiessling (ALMAnaCH)
arXiv Computer Vision
Aug 31

What Can Low Resource Languages Learn From Each Other?

The paper examines OCR adaptation for low‑resource languages, noting that fine‑tuning often hits a performance ceiling in data‑scarce settings. It identifies that lower layers of language‑specific models learn redundant features while higher layers capture script nuances, leading to a structural inefficiency. To address this, the authors propose PSMC, a framework that pre‑trains a base model, specializes it per language, merges the experts via task arithmetic, and co‑trains a unified multilingual backbone, achieving about a 2% improvement in Word Recognition Rate across 10 Indian scripts without adding parameters.

By Achyuth P, Kahaan Shah, Chetan Arora
arXiv Machine Learning
Aug 13

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

arXiv:2605. 16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography.

By Qinwu Xu, Yifan Jiang, Haoyu Ren
arXiv Machine Learning
Jul 9

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

arXiv:2607. 07179v1 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents.

By Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia
arXiv Computer Vision
Sep 22

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

The paper introduces an all‑in‑one multilingual scene text recognizer called ScriptMoE, which uses a script‑aware mixture‑of‑experts architecture to handle 10 scripts and 229 languages. It is built on a new large‑scale synthetic dataset, TextMuSS‑10M, and evaluated on the TextMuSS‑Bench, achieving 82.06% accuracy—1.31% higher than the best baseline. When integrated into the PP‑OCRv5 pipeline, ScriptMoE raises the end‑to‑end multilingual F1 score from 65.71% to 80.89%, slightly surpassing the best vision‑language model while using far fewer parameters.

By Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
arXiv Computation and Language
Aug 31

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

The paper introduces Synth-JDoc, a synthetic Japanese document image dataset created by rendering text with HTML and CSS to produce multi‑column layouts that include both vertical and horizontal writing styles. Images generated by text‑to‑image models are embedded to enhance visual realism, and noise and degradation filters are applied to improve robustness. Experiments show that fine‑tuning Large Vision Language Models on Synth‑JDoc yields superior performance on reading vertically written Japanese text compared to prior synthetic datasets.

By Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara