Mind the Gaps: A Curated Benchmark for Form Field Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2506. 14079v4 Announce Type: replace Abstract: Completing paperwork is a challenging and time-consuming problem.
arXiv:2607. 24745v1 Announce Type: cross Abstract: Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming manual process.
arXiv:2606. 17644v1 Announce Type: cross Abstract: Datasets in practical document processing scenarios typically grow over time, and their class annotations undergo continuous refinement.
Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur.
The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.
VietAIDetector is an open‑source, zero‑shot tool for detecting Vietnamese AI‑generated text. It offers a Gradio web interface that accepts raw Vietnamese text, common file formats, scanned documents, and very long texts beyond typical LLM context limits. Built on a Vietnamese‑specific language model, it outperforms existing English‑centric methods on out‑of‑domain datasets and lets users choose detection thresholds based on F1, accuracy, or TPR@0.05FPR, with results viewable or downloadable as a PDF report.