arXiv:2609.24220v1 Announce Type: new
Abstract: Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presen...
By Uday Allu (AI Research Team Yellow.ai), Abhivanth Sivaprakash (AI Research Team Yellow.ai), Pratik Singh (AI Research Team Yellow.ai), Aman Manocha (AI Research Team Yellow.ai)
arXiv:2610.02880v1 Announce Type: new
Abstract: Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this...
By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
arXiv:2608. 15064v1 Announce Type: new Abstract: Parsing visual documents into machine-readable representations is fundamental to document intelligence.
By Yuefeng Zou, Yichen Lu, Jingxiao Yang, Bingtao Fu, Gaoyang Zhang, Xiongfei Bai, Tian Chen, Xiang Qi
arXiv:2604.27724v2 Announce Type: replace
Abstract: Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich vis...
By Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang
arXiv:2607. 29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata.
By Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo
Jina-OCR-v1 is an end‑to‑end document parsing model designed for low‑budget GPUs, combining a compressed‑vision encoder with a 3B mixture‑of‑experts decoder that activates about 570 M parameters per token. It uses a FastMTP speculative decoding head that shares a single draft block across three prediction steps, with greedy verification ensuring lossless decoding. Post‑training includes instruction alignment, robustness fine‑tuning on difficult documents, and GRPO with dense verifiable rewards, achieving 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR‑Bench while delivering the highest page throughput at 2.57 pages per second on an NVIDIA L4 GPU.
By Alejandro Bar\'on Garc\'ia, Feng Wang, Emilia Garcia Casademont, Han Xiao
arXiv:2609.22628v1 Announce Type: new
Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We iso...
By Nikhil Reddy Pottanigari, Sepideh Kharaghani, Saverio Vadacchino, Alejandro Posada, Ying Zhang
arXiv:2603. 26815v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems for financial document QA typically follow a chunk-based paradigm: documents are split into fragments, embedded, and retrieved by similarity.
By Zhiyuan Cheng, Longying Lai, Yue Liu
arXiv:2608. 06146v1 Announce Type: new Abstract: End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence.
By Hao Yu, Jiabo Zhan, Kang Liu, Linnan Zhao, Dongxu Yue, Rui Chen, Jinglin Wang, Chong Sun, Chen Li, Jing Lyu, Chun Yuan
The paper introduces the Structural Semantic Unit (SSU) and the Coverage, Overlap, Trespass, and Excess (COTe) score as a new framework for evaluating Document Layout Analysis (DLA) models. Unlike traditional metrics such as IoU, F1, or mAP, which are tailored to 2D projections of 3D space, COTe focuses on the semantic structure of printed media and is decomposable to reveal specific failure modes like breaching semantic boundaries or redundant parsing. Experiments on five common DLA models across three datasets show that COTe is more informative and robust—especially under granularity mismatches—than F1, and the authors provide an SSU-labelled dataset and a Python library to facilitate adoption.
By Jonathan Bourne, Mwiza Simbeye, Ishtar Govia
SHELF is a Python system that creates synthetic, controlled benchmark data for evaluating large language models on bibliographic tasks such as classification, clustering, retrieval, pair classification, and instruction retrieval. It generates 62,899 model-written documents based on Library of Congress vocabularies and compares methods like TF, TF‑IDF, BM25, popular encoders, and zero‑shot decoders, reporting performance metrics such as 0.8887 for subject classification and 0.2605 for genre‑form classification. The tool also allows independent variation of bibliographic facets and can produce unseen documents beyond a model’s training cutoff, with results indicating that model rankings transfer more reliably than absolute scores when compared to other benchmarks.
By Michael J. Bommarito II
The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.
By Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst