FormGym: Doing Paperwork with Agents
arXiv:2506. 14079v4 Announce Type: replace Abstract: Completing paperwork is a challenging and time-consuming problem.
arXiv:2506. 14079v4 Announce Type: replace Abstract: Completing paperwork is a challenging and time-consuming problem.
arXiv:2607. 24745v1 Announce Type: cross Abstract: Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming manual process.
arXiv:2606. 17644v1 Announce Type: cross Abstract: Datasets in practical document processing scenarios typically grow over time, and their class annotations undergo continuous refinement.
Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur.
The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.
VietAIDetector is an open‑source, zero‑shot tool for detecting Vietnamese AI‑generated text. It offers a Gradio web interface that accepts raw Vietnamese text, common file formats, scanned documents, and very long texts beyond typical LLM context limits. Built on a Vietnamese‑specific language model, it outperforms existing English‑centric methods on out‑of‑domain datasets and lets users choose detection thresholds based on F1, accuracy, or TPR@0.05FPR, with results viewable or downloadable as a PDF report.
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
arXiv:2609.14352v1 Announce Type: new Abstract: AI-generated image detection has attracted increasing attention, but existing evaluations mainly focus on natural images, leaving AI-generated document...
arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...
arXiv:2609.01575v1 Announce Type: new Abstract: Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a...
arXiv:2609.08330v1 Announce Type: new Abstract: Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and vis...
arXiv:2606. 01393v1 Announce Type: cross Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems.