arXiv AI

Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models

arXiv:2607. 16201v1 Announce Type: new Abstract: Ontology engineering remains a critical bottleneck in knowledge-intensive AI systems.

arXiv AI
Aug 7

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

arXiv:2608. 06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard.

By Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem
arXiv Machine Learning
Jul 27

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

arXiv:2607. 21610v1 Announce Type: cross Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available.

By Miaobo Hu, Xiaobo Guo, Shuhao Hu, Bokun Wang, Rui Chen, Xin Wang, Daren Zha, Jun Xiao
arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
arXiv AI
Aug 11

H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System

arXiv:2608. 08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals.

By Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
arXiv AI
Jul 28

Retrieval-Augmented Generation of Ontologies from Relational Databases

arXiv:2506. 01232v2 Announce Type: replace-cross Abstract: Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph population, ontology-based data access, graph-based learning, and automated reasoning.

By Nadeen Fathallah, Mojtaba Nayyeri, Athish A Yogi, Ratan Bahadur Thapa, Hans-Michael Tautenhahn, Anton Schnurpel, Steffen Staab
arXiv Computation and Language
Sep 24

UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation

UniDataAgent (UniDataAgent) is an ontology‑grounded system designed to automate enterprise question‑to‑report tasks while preserving organization‑specific semantics. It separates semantic acquisition from online execution, with an Ontology Acquisition and Validation (OAV) stage that builds versioned ontologies from metadata, business knowledge, and expert input, and a Question‑to‑Report Execution (QRE) stage that retrieves semantic contracts, coordinates skills and data tools, validates results, and produces evidence‑linked reports. In a deployment across 27 enterprise tables and thousands of metric types, ontology construction took a few hours versus a week manually, and report generation took minutes versus several working days, achieving 95.0% strict accuracy on real business questions compared to 72.5% for document RAG.

By Yutai Duan, Yahui Zhao, Zhangti Li, Yu Ma, Zhenfeng Qi, Shaoyang Yuan, Jing Fan, Jie Liu