arXiv:2609.37195v1 Announce Type: new
Abstract: Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down ti...
By Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza
ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.
By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
By Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata
The paper adapts the compact PP‑OCRv6 recognizer for historical text recognition and compares it to a conventional CRNN across various training regimes, including generalized pretraining, domain‑specific training, corpus‑level fine‑tuning, and manuscript‑specific few‑shot adaptation on multilingual Latin and Arabic scripts. While PP‑OCRv6 does not always beat the CRNN when trained from scratch, heterogeneous pretraining significantly improves its generalization. Additionally, fine‑tuned PP‑OCRv6 can surpass a large vision‑language model (Qwen3.5‑based Medusa) that is specifically tailored for historical Latin‑script handwriting recognition.
By Benjamin Kiessling (ALMAnaCH)
arXiv:2609.08188v1 Announce Type: new
Abstract: Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retriever...
By Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour, Mung Chiang, Christopher G. Brinton
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv:2606. 29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets.
By Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored.
arXiv:2607. 03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details.
By Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu
arXiv:2607. 07179v1 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents.
By Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv:2509. 07295v4 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture.
By Ji Xie, Trevor Darrell, Luke Zettlemoyer, XuDong Wang