arXiv AI By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid

Smart Content Ingestion for Generative AI Workloads

Read the original on arXiv AI →

The paper introduces a production-ready content‑extraction system tailored for generative AI workloads, addressing the heterogeneity of enterprise data formats such as PDFs, spreadsheets, and scanned documents. It features selective OCR routing, a scarcity‑first curation engine with a reference‑based extraction scorer, a deterministic structure‑aware chunker, and a read‑only retrieval evaluator that generates grounded questions and reports metrics like Hit@k and MRR. On a 180‑document corpus, the system achieves high accuracy (97.4/100 character score, 0.13% error rate) and strong retrieval performance (Hit@1 68.6%, Hit@10 92.8%, MRR 0.77).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
arXiv Computer Vision
Sep 22

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

arXiv:2609.24220v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presen...

By Uday Allu (AI Research Team Yellow.ai), Abhivanth Sivaprakash (AI Research Team Yellow.ai), Pratik Singh (AI Research Team Yellow.ai), Aman Manocha (AI Research Team Yellow.ai)
arXiv AI
Sep 12

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.

By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou