arXiv AI

ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data

arXiv:2607. 18262v1 Announce Type: new Abstract: The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products.

arXiv Computation and Language
Sep 1

RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation

RENSA is a federated SPARQL query generation framework that extends SPARQL Builder Metadata to include class and authority information, enabling precise source selection and semantic constraint inference without runtime ASK queries. The generated metadata profiles occupy less than 1% of the original dataset triples, providing storage‑efficient insights. Evaluation on the LargeRDFBench benchmark shows that RENSA matches state‑of‑the‑art source selection performance while eliminating runtime communication overhead.

By Victor Eiti Yamamoto, Takeda Hideaki, Yamamoto Yasunori
arXiv AI
Aug 28

From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement

The paper introduces a novel LLM‑driven multi‑agent pipeline that converts relational databases into graph databases by standardizing table and column names and iteratively refining the graph schema through ETL, Analyzer, and Graph agents. The resulting graph database meets accuracy, groundedness, and faithfulness criteria and shows significant performance gains, achieving 85.6% Q&A accuracy—12.12% higher than an SQL agent on PostgreSQL—and reducing latency by roughly threefold on a BFSI dataset. This demonstrates an efficient, automated method for transforming tabular data into a more intuitive and faster‑executing graph format.

By Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy
arXiv AI
Jun 8

MetaConfigurator: AI-Assisted RDF Authoring from JSON Data

arXiv:2606. 07094v1 Announce Type: cross Abstract: Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across systems due to lacking semantic interoperability.

By Felix Neubauer, Mahdi Jafarkhani, Kenichi Endo, J\"urgen Pleiss, Benjamin Uekermann
arXiv AI
Aug 17

From Field Data to Global Food Systems Intelligence: A Semantic Graph Framework for Sustainable Wheat Production

arXiv:2502. 19507v2 Announce Type: replace Abstract: In response to the growing need for structured, interoperable agricultural data, this paper presents the Sustainable Wheat Production Datahub, a modular, graph-based framework that brings diverse wheat production datasets together into a single, queryable store.

By Nirmal Gelal, Aastha Gautam, Soheil Abadifard, Nico Giordano, Moumita Sen Sarma, Sanaz Saki Norouzi, Claudio Dias da Silva Jr, Jean Ribert Francois, Kathleen M. Jagodnik, Katherine Nelson, Terry Griffin, Xiaomao Lin, Stacy Hutchinson, Stephen M. Welch, Kelsey Andersen Onofre, Romulo Lollato, Pascal Hitzler, Hande K\"u\c{c}\"uk McGinty
arXiv AI
Jul 28

Retrieval-Augmented Generation of Ontologies from Relational Databases

arXiv:2506. 01232v2 Announce Type: replace-cross Abstract: Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph population, ontology-based data access, graph-based learning, and automated reasoning.

By Nadeen Fathallah, Mojtaba Nayyeri, Athish A Yogi, Ratan Bahadur Thapa, Hans-Michael Tautenhahn, Anton Schnurpel, Steffen Staab
arXiv Machine Learning
Sep 11

Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)

The paper introduces EXYGEN, a framework that enables conversational access to large knowledge graphs by combining VoID descriptions, ShEx schemas, retrieved triples, and example question‑query pairs in a retrieval‑augmented generation pipeline. On the SciQA benchmark, this approach achieves an exact‑match score of 0.419 without fine‑tuning any large language model, and shows that larger general‑purpose LLMs can outperform smaller code‑specialized ones when provided sufficient context. To scale metadata generation for very large KGs, the authors propose a predicate‑coverage‑aware parallel graph sampling strategy that preserves structural diversity, reduces runtime by over 80× on OpenCitations Meta and GESIS, and is the only tractable method for obtaining complete metadata on ORKG.

By Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello
arXiv AI
2d ago

Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying

Build2SPARQL is a large-scale benchmark dataset for translating natural-language questions into SPARQL queries over building knowledge graphs. The dataset is generated by a KG‑grounded pipeline that produces 6,136 executable SPARQL queries and 30,680 corresponding natural-language questions across six query-pattern families and five vocabulary registers, covering 201 building KGs. Human validation shows high semantic fidelity, naturalness, and operational plausibility, and retrieval‑augmented evaluation demonstrates significant accuracy gains for open‑weight language models.

By Wooyoung Jung