arXiv AI

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.

arXiv AI
Aug 20

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

The paper evaluates open-source OCR, LLM, and VLM systems on a high‑risk public sector task: extracting structured data from student application documents. Results show that VLMs generally outperform OCR+LLM pipelines, yet only 4 of 35 configurations achieve F1 scores above 0.5, with most combinations scoring below 0.25. Model size and input quality, especially preserving OCR structure, are critical factors influencing performance.

By Elias Schubert, Felix Bie{\ss}mann
arXiv Computer Vision
Sep 18

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HunyuanOCR-1.5 is a lightweight, end‑to‑end OCR‑specialized vision‑language model that unifies document parsing, text spotting, information extraction, text‑image translation, and multi‑image document understanding. It builds on the HunyuanOCR‑1.0 architecture, improving efficiency with DFlash‑based OCR decoding for faster inference (6.37× Transformer speedup, 2.14× under vLLM) and enhancing capability through an Agentic Data Flow system that autonomously constructs high‑quality training data for long‑tail OCR tasks. The model achieves top‑tier performance on OmniDocBench v1.6 and sets new milestones in ancient‑script OCR, chart/table parsing, multilingual parsing, and hallucination evaluation, while remaining lightweight for deployment.

By Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, Yu Zhou
arXiv Computation and Language
Sep 4

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

Jina-OCR-v1 is an end‑to‑end document parsing model designed for low‑budget GPUs, combining a compressed‑vision encoder with a 3B mixture‑of‑experts decoder that activates about 570 M parameters per token. It uses a FastMTP speculative decoding head that shares a single draft block across three prediction steps, with greedy verification ensuring lossless decoding. Post‑training includes instruction alignment, robustness fine‑tuning on difficult documents, and GRPO with dense verifiable rewards, achieving 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR‑Bench while delivering the highest page throughput at 2.57 pages per second on an NVIDIA L4 GPU.

By Alejandro Bar\'on Garc\'ia, Feng Wang, Emilia Garcia Casademont, Han Xiao
Hugging Face Trending Papers
Aug 6

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context.

arXiv Computation and Language
Sep 2

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

arXiv:2609.01575v1 Announce Type: new Abstract: Extracting structured fields from hundreds of millions of documents annually remains costly in regulated industries: bespoke OCR cascades cover only a...

By Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin, Olga Tsymboi, Anatolii Potapov, Aleksandr Ivanov
arXiv AI
Sep 15

Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction

The paper evaluates eleven vision‑language models (VLMs) for extracting structured fields from business documents, focusing on robustness, cost, and governance rather than just accuracy. Using a held‑out set of 750 synthetic checks, the study finds that fine‑tuning open‑source VLMs on 3,000 samples yields an F1 score above 0.98, surpassing all zero‑shot commercial systems, while GPT‑5 tops the commercial group and Claude Sonnet 4.5 fails on date extraction. The authors also present a practitioner‑oriented selection framework that maps task profiles—such as quality, latency, governance, and volume—to recommended approaches via filtering and total‑cost minimization, demonstrated on a mid‑volume document‑extraction scenario.

By Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li, Junlin Jiang
arXiv Computation and Language
Aug 24

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

arXiv:2608.20868v1 Announce Type: cross Abstract: Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information ext...

By A. Said Gurbuz (IBM Research Zurich, ETH Zurich), Ahmed Nassar (IBM Research Zurich), Christoph Auer (IBM Research Zurich), Maksym Lysak (IBM Research Zurich), Lucas Morin (IBM Research Zurich), Matteo Omenetti (IBM Research Zurich), Tim Strohmeyer (IBM Research Zurich), Panagiotis Vagenas (IBM Research Zurich), Nikolaos Livathinos (IBM Research Zurich), Michele Dolfi (IBM Research Zurich), Peter Staar (IBM Research Zurich)
Hugging Face Trending Papers
Jul 21

RAGAL: A Frugal, Fully Local Retrieval-Augmented Assistant for Technical Support at a Government Agency

Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine.

arXiv Machine Learning
Jun 30

DataComp-VLM: Improved Open Datasets for Vision-Language Models

arXiv:2606. 28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies.

By Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan Hammoud, Thomas De Min, Simone Caldarella, Jehanzeb Mirza, Sedrick Keh, Mehdi Cherti, Hilde Kuehne, Bernt Schiele, Serena Yeung-Levy, Muhammad Ferjad Naeem, Federico Tombari, Ana Klimovic, Elisa Ricci, Matthias Bethge, Sewoong Oh, Ameya Prabhu, Alessio Tonioni, Jenia Jitsev, Massimiliano Mancini, Ludwig Schmidt, Nikhil Parthasarathy