Towards Data Science

One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited

Enterprise Document Intelligence [Vol. 1 #9B] - One call wires the four upgraded bricks together, run on a paper, a NIST standard, and a report with a broken TOC The post One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited appeared first on Towards Data Science .

Towards Data Science
Jun 9

10 Common RAG Mistakes We Keep Seeing in Production

Enterprise Document Intelligence [Vol. 1 #4bis] - A coauthor note on the brick-by-brick pitfalls that justified the four-brick split, before Part II walks the fixes The post 10 Common RAG Mistakes We Keep Seeing in Production appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Sep 2

A RAG That Says “Not in This Document” Has to Show Four Kinds of Evidence

The article discusses the importance of a Retrieval-Augmented Generation (RAG) system providing clear evidence when it states that information is not present in a document. It outlines four distinct types of evidence that should accompany such a claim to avoid presenting a confident but incorrect answer or an unsupported “no answer.” The piece emphasizes that each evidence type serves as a safeguard against misinformation in enterprise document intelligence.

By Kezhan Shi
Towards Data Science
Aug 22

Multi-Document RAG: A Folder of Unrelated PDFs Is One Long Document with a Nested Outline

The article discusses a method for handling a folder of unrelated PDFs as a single long document with a nested outline. It highlights that without shared fields, an index cannot be built, so the approach uses one summary line per file and each file’s own table of contents, with retrieval routes extending down two levels. This structure enables retrieval-augmented generation (RAG) across multiple documents.

By angela shi
Towards Data Science
Aug 20

Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One

The article explains that enterprise document intelligence can be categorized into three distinct corpus types, each requiring a specific architecture. It outlines how to determine the shape of a document collection through three key questions. The piece also discusses the costs associated with building a system for the incorrect corpus type.

By angela shi