← Back to all news
Towards Data Science June 12, 2026 By Kezhan Shi

When PyMuPDF Can’t See the Table: Parse PDFs for RAG with Azure Layout

Read the original on Towards Data Science →

Enterprise Document Intelligence [Vol. 1 #5bis] - The same relational tables.

Summary generated by The Flow from the publisher's feed. The full article lives at Towards Data Science.

  • rag
  • computer-vision

Related stories

Towards Data Science
Jun 13

Parse PDFs for RAG Locally with Docling: Rich Tables, No Cloud Upload

Enterprise Document Intelligence [Vol. 1 #5ter] - Table cells, OCR, captions, headings: cloud-grade structure, running on your own machine.

By Kezhan Shi
ragcomputer-vision
More like this →
Towards Data Science
Jun 11

Stop Returning Flat Text from a PDF: The Relational Tables RAG Needs

Enterprise Document Intelligence [Vol. 1 #5B] - One PDF in, a relational set of DataFrames out: lines, pages, TOC, images, cross-references, captions, spans, and a parsing summary The post Stop Returning Flat Text from a PDF: The Relational Tables RAG Needs appeared first on Towards Data Science .

By Kezhan Shi
rag
More like this →
Towards Data Science
Jun 11

Stop Returning Flat Text from a PDF: The Relational Shape RAG Needs

Enterprise Document Intelligence [Vol. 1 #5B] - One PDF in, a relational set of DataFrames out: lines, pages, TOC, images, cross-references, captions, spans, and a parsing summary The post Stop Returning Flat Text from a PDF: The Relational Shape RAG Needs appeared first on Towards Data Science .

By Kezhan Shi
rag
More like this →
Towards Data Science
Jun 10

Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality

Enterprise Document Intelligence [Vol. 1 #5A] - Document signals (metadata, native TOC, source software) and page-level content (text vs scans, tables, images, columns, page profile) The post Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality appeared first on Towards Data Science .

By Kezhan Shi
rag
More like this →
Towards Data Science
Jul 7

A Production RAG Pipeline for PDFs: Relational Parsing, TOC Retrieval, Typed Answers

Enterprise Document Intelligence [Vol. 1 #9A] - Same paper, same question as Article 1.

By angela shi
rag
More like this →
Towards Data Science
Jul 17

One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited

Enterprise Document Intelligence [Vol. 1 #9B] - One call wires the four upgraded bricks together, run on a paper, a NIST standard, and a report with a broken TOC The post One RAG Pipeline, Four Very Different PDFs: Same Four Bricks, Every Answer Typed and Cited appeared first on Towards Data Science .

By angela shi
rag
More like this →