Enterprise Document Intelligence [Vol. 1 #5A] - Document signals (metadata, native TOC, source software) and page-level content (text vs scans, tables, images, columns, page profile) The post Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality appeared first on Towards Data Science .
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5ter] - Table cells, OCR, captions, headings: cloud-grade structure, running on your own machine.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5nonies] - Nature, plan, execute, synthesize: closing brick 1 with a dispatcher that reads each PDF’s nature and picks the method that fits, fitz, Docling, PaddleOCR, EasyOCR, MinerU or Surya, then folds the outputs into one corpus The post Before Full Agentic RAG: Know How You Decide, and the Parsing Methods You Pick From appeared first on Towards Data Science .
By angela shi
Enterprise Document Intelligence [Vol. 1 #5sexies] - image_df tells you where every picture is.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #1] The smallest version of RAG that actually works, on a real PDF, with grounded answers and the source lines highlighted.
By angela shi
Enterprise Document Intelligence [Vol. 1 #5B] - One PDF in, a relational set of DataFrames out: lines, pages, TOC, images, cross-references, captions, spans, and a parsing summary The post Stop Returning Flat Text from a PDF: The Relational Shape RAG Needs appeared first on Towards Data Science .
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5quater] - The other parsers read the words on a page.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #9A] - Same paper, same question as Article 1.
By angela shi
The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5octies] - Rules propose, LLM validates: six deterministic signals on span-level typography surface heading candidates, one bounded loop keeps the real ones, and the same toc_df drops back into the RAG pipeline The post Building Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG appeared first on Towards Data Science .
By angela shi
The article discusses how noisy text—stemming from user typos, rapid transcription errors, and OCR character mistakes—poses challenges for Retrieval-Augmented Generation (RAG) systems. It explains that traditional spell-checking only addresses one type of error, while embeddings are needed to handle the remaining noise. The piece highlights the need for more robust solutions in enterprise document intelligence.
By Kezhan Shi
Enterprise Document Intelligence [Vol. 1 #5B] - One PDF in, a relational set of DataFrames out: lines, pages, TOC, images, cross-references, captions, spans, and a parsing summary The post Stop Returning Flat Text from a PDF: The Relational Tables RAG Needs appeared first on Towards Data Science .
By Kezhan Shi