arXiv AI By AJ Carl P. Dy, Aivin V. Solatorio

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Read the original on arXiv AI →

arXiv:2606. 06242v1 Announce Type: cross Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 4

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts.

arXiv AI
Jul 17

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

arXiv:2607. 14117v1 Announce Type: cross Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous elements such as text, tables, formulas, figures, and layout cues.

By Zhen Yina, Wenkang An, Hao Wang, Keran You