The article explains that enterprise document intelligence can be categorized into three distinct corpus types, each requiring a specific architecture. It outlines how to determine the shape of a document collection through three key questions. The piece also discusses the costs associated with building a system for the incorrect corpus type.
By angela shi
Enterprise Document Intelligence [Vol. 1 #9B] - One call wires the four upgraded bricks together, run on a paper, a NIST standard, and a report with a broken TOC The post A Production RAG Pipeline in Action: Every Answer Typed and Cited appeared first on Towards Data Science .
By angela shi
Enterprise Document Intelligence [Vol. 1 #9B] - One call wires the four upgraded bricks together, run on a paper, a NIST standard, and a report with a broken TOC The post One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited appeared first on Towards Data Science .
By angela shi
Enterprise Document Intelligence [Vol. 1 #7ter] - Six positions on the retrieval brick that contradict the cosine-first reflex of mainstream RAG The post The Untaught Lessons of RAG Retrieval: Cosine Is Not the Foundation appeared first on Towards Data Science .
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5sexies] - image_df tells you where every picture is.
By Kezhan Shi
The true bottleneck was never the analysis. The post BI Is Dead, Long Live BI appeared first on Towards Data Science .
By Mahdi Karabiben
Enterprise Document Intelligence [Vol. 1 #7C] - One LLM call ranks the candidates with reasons.
By angela shi
Enterprise Document Intelligence [Vol. 1 #4] - A diagnostic across PDFs and questions, and a map of the techniques the rest of the series will cover The post From Regex to Vision Models: Which RAG Technique Fits Which Problem appeared first on Towards Data Science .
By angela shi
The article discusses how noisy text—stemming from user typos, rapid transcription errors, and OCR character mistakes—poses challenges for Retrieval-Augmented Generation (RAG) systems. It explains that traditional spell-checking only addresses one type of error, while embeddings are needed to handle the remaining noise. The piece highlights the need for more robust solutions in enterprise document intelligence.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #4bis] - A coauthor note on the brick-by-brick pitfalls that justified the four-brick split, before Part II walks the fixes The post 10 Common RAG Mistakes We Keep Seeing in Production appeared first on Towards Data Science .
By Kezhan Shi
The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5A] - Document signals (metadata, native TOC, source software) and page-level content (text vs scans, tables, images, columns, page profile) The post Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality appeared first on Towards Data Science .
By Kezhan Shi