arXiv AI

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

WildHandBench is a new benchmark comprising 500 handwritten documents that span free text, tables, and formulas across four languages and nine real‑world scenarios. It introduces a Prior‑Driven Error (PDE) metric to assess whether mistakes stem from language priors rather than visual cues. In tests of 18 state‑of‑the‑art models, the best achieves only 71.85% accuracy, while humans reach 77.09%, and model errors are largely prior‑driven (63‑91%) compared to human errors (49%).

arXiv AI
Aug 20

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

OmniHandwritingOCR is a diagnostic benchmark designed to evaluate multimodal large language models (MLLMs) and OCR systems on handwritten text and mathematical expression recognition. It comprises 77.57K labeled images across six subtasks and twelve subsets, including a difficulty‑stratified multi‑line formula corpus that tests robustness to increasing structural complexity. The benchmark reveals that current systems perform poorly on complex multi‑line formulas, exhibit variable rankings across languages and formula settings, and sometimes hallucinate corrections that are not visually supported.

By Zinuo Guo, Min Zhang, Bo Jiang
Hugging Face Trending Papers
Jun 24

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task.

arXiv AI
Jun 2

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

arXiv:2606. 01393v1 Announce Type: cross Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems.

By Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo, Zhenting Qi, Konwoo Kim, Longtian Ye, Xiaolong Luo, Jinhe Bi, Henry Zhang, Haris Riaz, Xuan Zhang, Yunze Xiao, Bangya Liu, Tom Tang, Yunfei Zhao, Qunshu Lin, Zihan Wang, Minghao Liu, Michael Lingzhi Li, Yilun Du, Jesse Thomason, Rogerio Feris, Alex Pentland, Zexue He
arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv AI
Jun 12

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

arXiv:2602. 00122v3 Announce Type: replace-cross Abstract: In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language instructions.

By Hongzhu Yi, Yujia Yang, Yuanxiang Wang, Tong Li, Zhenyu Guan, Tianyu Zong, Jiahuan Chen, Chenxi Bao, Tiankun Yang, Haopeng Jin, Yixuan Yuan, Xinming Wang, Tao Yu, Ruilin Gao, Ruiwen Tao, Haijin Liang, Jin Ma, Jinwen Luo, Yeshani, Xinyu Zuo, Jungang Xu
arXiv Machine Learning
Sep 14

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

ExpertHTR is a unified vision‑language framework for handwritten text recognition that tackles the challenge of small, heterogeneous datasets by organizing structural annotations into a common Page‑Region‑Line representation. It defines four related training tasks—complete transcription, physical‑line coverage, text localization, and localized recognition—without extra manual labels. The model combines a jointly trained dense backbone with a sparse Mixture‑of‑Experts architecture, using Sparsegen routing and regularization to adaptively activate experts, achieving state‑of‑the‑art results on the IAM benchmark and outperforming general‑purpose OCR systems on most datasets.

By Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy