arXiv:2609.23679v1 Announce Type: new
Abstract: Form Field Detection (FFD) is a fundamental component of document understanding systems, enabling applications ranging from large-scale industrial digi...
By Iheb Brini, Omar Moured, Hamza Gbada, Elisa Barney
The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.
By Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li, Junlin Jiang
arXiv:2609.22628v1 Announce Type: new
Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We iso...
By Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino, Alejandro Posada, Ying Zhang
arXiv:2608. 07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts.
By Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining Guo
We present HunyuanOCR-1. 5, a lightweight end-to-end OCR-specialized vision-language model.
arXiv:2502. 20295v3 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task.
By Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong