Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2604.27724v2 Announce Type: replace Abstract: Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich vis...
The paper investigates how general vision‑language models (VLMs) develop specialized optical character recognition (OCR) capabilities. By applying a causal intervention protocol, the authors identify sparse, stable OCR‑head sets in several VLMs and show that these heads largely overlap with textual retrieval/copy heads found in general VLMs. The study concludes that full‑sequence OCR functions as a dense multimodal copy‑and‑paste mechanism, and that when a VLM is fine‑tuned for OCR, it largely preserves the same head identities while redistributing their functional and causal strengths.
arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We iso...
arXiv:2605.27243v3 Announce Type: replace Abstract: Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent traject...
The paper introduces a lightweight CPU-based extension to GROBID that uses layout-guided masking to identify figure, table, and paratext regions in scientific PDFs. By routing tokens to specialized GROBID models or discarding them, the method improves structural accuracy on PMC corpora and enhances figure caption recovery. It also achieves competitive table detection and body‑text precision compared to vision‑based GPU parsers while operating entirely on CPU and costing significantly less.
The paper introduces Invoice Haystack, a benchmark of 1,500 anonymized invoices and 200 question‑answer pairs that tests document retrieval and visual question answering under strong visual homogeneity. It shows that existing benchmarks suffer from embedding collapse, with Invoice Haystack’s mean pairwise cosine similarity at 0.73 versus 0.38 and 0.31 in DocHaystack and InfoHaystack. The authors propose VL‑RAG, a hybrid retrieval‑augmented generation framework that combines text and visual embeddings and a VLM‑based verification filter, achieving 60.0% Recall@1 on Invoice Haystack‑500 and improving performance on other benchmarks.