Towards Data Science

Reconstructing the Table of Contents a PDF Forgot to Ship, So RAG Can Scope by Section

Enterprise Document Intelligence [Vol. 1 #5septies] - When a PDF prints a contents page but exposes no outline, two ways to turn it back into structure, plus the page-alignment step everyone forgets The post Reconstructing the Table of Contents a PDF Forgot to Ship, So RAG Can Scope by Section appeared first on Towards Data Science .

Towards Data Science
Aug 22

Multi-Document RAG: A Folder of Unrelated PDFs Is One Long Document with a Nested Outline

The article discusses a method for handling a folder of unrelated PDFs as a single long document with a nested outline. It highlights that without shared fields, an index cannot be built, so the approach uses one summary line per file and each file’s own table of contents, with retrieval routes extending down two levels. This structure enables retrieval-augmented generation (RAG) across multiple documents.

By angela shi
Towards Data Science
Sep 3

Tables in PDFs for RAG: Don’t Flatten the Grid

The article titled "Tables in PDFs for RAG: Don’t Flatten the Grid" discusses Enterprise Document Intelligence, presenting a diagnostic and five composable operations rather than a decision tree. It appears in the Vol.1 #B4 issue of the publication and was first posted on Towards Data Science.

By Kezhan Shi
Towards Data Science
Aug 5

Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG

Enterprise Document Intelligence [Vol. 1 #5octies] - Rules propose, LLM validates: six deterministic signals on span-level typography surface heading candidates, one bounded loop keeps the real ones, and the same toc_df drops back into the RAG pipeline The post Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG appeared first on Towards Data Science .

By angela shi
Towards Data Science
Aug 23

Parse the Folder, Not Just the PDFs: The Relational Tables RAG Needs on a Case File

The article discusses how Enterprise Document Intelligence should begin by parsing the folder structure rather than just PDFs, emphasizing that the index must reflect the case type’s requirements before any folder is accessed. It highlights that the two key questions to develop are not retrieval questions but rather focus on the relational tables needed for Retrieval-Augmented Generation (RAG) in a case file context.

By angela shi