Towards Data Science

Parse Scanned PDFs for RAG with EasyOCR: Free OCR Gives You Words, Not a Document

Enterprise Document Intelligence [Vol. 1 #5quinquies] - Same 1974 scanned PDF, two engines.

Towards Data Science
Aug 12

Before Full Agentic RAG: Know How You Decide, and the Parsing Methods You Pick From

Enterprise Document Intelligence [Vol. 1 #5nonies] - Nature, plan, execute, synthesize: closing brick 1 with a dispatcher that reads each PDF’s nature and picks the method that fits, fitz, Docling, PaddleOCR, EasyOCR, MinerU or Surya, then folds the outputs into one corpus The post Before Full Agentic RAG: Know How You Decide, and the Parsing Methods You Pick From appeared first on Towards Data Science .

By angela shi
Towards Data Science
Aug 29

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.

By Kezhan Shi
Towards Data Science
Aug 5

Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG

Enterprise Document Intelligence [Vol. 1 #5octies] - Rules propose, LLM validates: six deterministic signals on span-level typography surface heading candidates, one bounded loop keeps the real ones, and the same toc_df drops back into the RAG pipeline The post Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG appeared first on Towards Data Science .

By angela shi
Towards Data Science
Aug 30

Noisy Text in RAG: Typos, OCR, and the Gap Classical Spell-Check Leaves

The article discusses how noisy text—stemming from user typos, rapid transcription errors, and OCR character mistakes—poses challenges for Retrieval-Augmented Generation (RAG) systems. It explains that traditional spell-checking only addresses one type of error, while embeddings are needed to handle the remaining noise. The piece highlights the need for more robust solutions in enterprise document intelligence.

By Kezhan Shi