arXiv Computation and Language By Luca Foppiano, Sana Khamassi, Vipul Gupta

Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion

Read the original on arXiv Computation and Language →

The paper introduces a lightweight CPU-based extension to GROBID that uses layout-guided masking to identify figure, table, and paratext regions in scientific PDFs. By routing tokens to specialized GROBID models or discarding them, the method improves structural accuracy on PMC corpora and enhances figure caption recovery. It also achieves competitive table detection and body‑text precision compared to vision‑based GPU parsers while operating entirely on CPU and costing significantly less.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computer Vision
Sep 22

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

arXiv:2609.24220v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presen...

By Uday Allu (AI Research Team Yellow.ai), Abhivanth Sivaprakash (AI Research Team Yellow.ai), Pratik Singh (AI Research Team Yellow.ai), Aman Manocha (AI Research Team Yellow.ai)
arXiv Computation and Language
Sep 4

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

Jina-OCR-v1 is an end‑to‑end document parsing model designed for low‑budget GPUs, combining a compressed‑vision encoder with a 3B mixture‑of‑experts decoder that activates about 570 M parameters per token. It uses a FastMTP speculative decoding head that shares a single draft block across three prediction steps, with greedy verification ensuring lossless decoding. Post‑training includes instruction alignment, robustness fine‑tuning on difficult documents, and GRPO with dense verifiable rewards, achieving 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR‑Bench while delivering the highest page throughput at 2.57 pages per second on an NVIDIA L4 GPU.

By Alejandro Bar\'on Garc\'ia, Feng Wang, Emilia Garcia Casademont, Han Xiao