The paper examines OCR adaptation for low‑resource languages, noting that fine‑tuning often hits a performance ceiling in data‑scarce settings. It identifies that lower layers of language‑specific models learn redundant features while higher layers capture script nuances, leading to a structural inefficiency. To address this, the authors propose PSMC, a framework that pre‑trains a base model, specializes it per language, merges the experts via task arithmetic, and co‑trains a unified multilingual backbone, achieving about a 2% improvement in Word Recognition Rate across 10 Indian scripts without adding parameters.
By Achyuth P, Kahaan Shah, Chetan Arora
The paper introduces PuMVR, a benchmark of 1,000 Punjabi image‑text pairs spanning three scripts—Gurmukhi, Shahmukhi, and Roman—to evaluate Vision‑Language Models (VLMs). Testing ten state‑of‑the‑art VLMs reveals a significant Script Gap: models perform well in one script but poorly in another, with accuracy differences up to 16%. The authors propose the Script Consistency Rate (SCR) as a new metric, noting it can be as low as 24.8% on their benchmark, and argue that current multilingual VLMs are not truly multi‑script.
By Prabhjot Singh, Bhushan Pawar, Madhu Reddiboina, Rajvee Sheth
arXiv:2605. 16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography.
By Qinwu Xu, Yifan Jiang, Haoyu Ren
ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.
By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
The paper introduces PM4Bench, a multimodal, multilingual, multi-task benchmark built on a strictly parallel 10‑language corpus, allowing fair cross‑lingual comparison of Large Vision‑Language Models (LVLMs). It also proposes a vision setting that embeds textual inputs directly into images to better mimic real deployment scenarios. Experiments show OCR performance drives cross‑lingual gaps, leading to an OCR‑centric GRPO training strategy that improves multilingual VQA and reduces disparities without costly supervision.
By Junyuan Gao, Jiahe Song, Jiang Wu, Runchuan Zhu, Guanlin Shen, Shasha Wang, Xingjian Wei, Haote Yang, Weijia Li, Bin Wang, Lijun Wu, Conghui He
MEVL-STP introduces a two‑stage pipeline for spotting arbitrarily shaped scene text. The detection stage fuses features from six frozen vision encoders via a hierarchical Feature Pyramid Network and a Progressive Scale Expansion network to produce precise polygon masks. The recognition stage then crops these masks and feeds them to a fine‑tuned Qwen3‑VL‑8B‑Instruct model, achieving state‑of‑the‑art detection and end‑to‑end performance on CTW1500, Total‑Text, and ICDAR 2015 without synthetic pretraining.
By Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote