Towards Data Science

Most RAG Hallucinations Are Retrieval Failures: How the Retrieval Brick Decides What the Model Can Invent

Enterprise Document Intelligence [Vol. 1 #7quinquies] - Hallucination is usually garbage-in.

Towards Data Science
Jun 9

10 Common RAG Mistakes We Keep Seeing in Production

Enterprise Document Intelligence [Vol. 1 #4bis] - A coauthor note on the brick-by-brick pitfalls that justified the four-brick split, before Part II walks the fixes The post 10 Common RAG Mistakes We Keep Seeing in Production appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Sep 2

A RAG That Says “Not in This Document” Has to Show Four Kinds of Evidence

The article discusses the importance of a Retrieval-Augmented Generation (RAG) system providing clear evidence when it states that information is not present in a document. It outlines four distinct types of evidence that should accompany such a claim to avoid presenting a confident but incorrect answer or an unsupported “no answer.” The piece emphasizes that each evidence type serves as a safeguard against misinformation in enterprise document intelligence.

By Kezhan Shi
Towards Data Science
Aug 29

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.

By Kezhan Shi