arXiv AI

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

arXiv:2606. 06242v1 Announce Type: cross Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables.

Hugging Face Trending Papers
Jun 4

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts.

arXiv Computer Vision
Sep 18

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large‑scale benchmark for fine‑grained, element‑level source attribution in Document Visual Question Answering (VQA). It introduces MAPPET, a Mask‑based Perplexity‑Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question‑answer pairs with element‑level grounding, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality, revealing that even strong models often fail to localize supporting elements.

By Luca De Grandis (University of Modena and Reggio Emilia, Modena, Italy), Silvia Cappelletti (University of Modena and Reggio Emilia, Modena, Italy), William Raccagni (University of Modena and Reggio Emilia, Modena, Italy, University of Pisa, Pisa, Italy), Marcella Cornia (University of Modena and Reggio Emilia, Modena, Italy), Lorenzo Baraldi (University of Modena and Reggio Emilia, Modena, Italy), Rita Cucchiara (University of Modena and Reggio Emilia, Modena, Italy)
Hugging Face Trending Papers
Sep 17

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large-scale benchmark that provides fine-grained, element-level source attribution for Document Visual Question Answering (VQA). It introduces MAPPET, a Mask-based Perplexity-Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question-answer pairs with grounding annotations, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality.

arXiv Computer Vision
2d ago

The COTe score: A decomposable framework for evaluating Document Layout Analysis models

The paper introduces the Structural Semantic Unit (SSU) and the Coverage, Overlap, Trespass, and Excess (COTe) score as a new framework for evaluating Document Layout Analysis (DLA) models. Unlike traditional metrics such as IoU, F1, or mAP, which are tailored to 2D projections of 3D space, COTe focuses on the semantic structure of printed media and is decomposable to reveal specific failure modes like breaching semantic boundaries or redundant parsing. Experiments on five common DLA models across three datasets show that COTe is more informative and robust—especially under granularity mismatches—than F1, and the authors provide an SSU-labelled dataset and a Python library to facilitate adoption.

By Jonathan Bourne, Mwiza Simbeye, Ishtar Govia
arXiv AI
Jul 17

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

arXiv:2607. 14117v1 Announce Type: cross Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous elements such as text, tables, formulas, figures, and layout cues.

By Zhen Yina, Wenkang An, Hao Wang, Keran You
arXiv AI
Jul 21

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.

By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv AI
Jun 4

MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A

arXiv:2606. 04231v1 Announce Type: cross Abstract: Recent advances in multimodal retrieval-augmented generation (MM-RAG) have shifted toward minimal parsing, relying on page-level images for producing retriever embeddings and for answer generation.

By Hanoz Bhathena, Parin Rajesh Jhaveri, Rohan Mittal, Prateek Singh, Aymen Kallala, Rachneet Kaur, Yiqiao Jin, Zhen Zeng, Adwait Ratnaparkhi, Denis Kochedykov