arXiv Machine Learning

AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs

arXiv:2608. 14228v1 Announce Type: new Abstract: Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links.

arXiv Machine Learning
Sep 11

Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)

The paper introduces EXYGEN, a framework that enables conversational access to large knowledge graphs by combining VoID descriptions, ShEx schemas, retrieved triples, and example question‑query pairs in a retrieval‑augmented generation pipeline. On the SciQA benchmark, this approach achieves an exact‑match score of 0.419 without fine‑tuning any large language model, and shows that larger general‑purpose LLMs can outperform smaller code‑specialized ones when provided sufficient context. To scale metadata generation for very large KGs, the authors propose a predicate‑coverage‑aware parallel graph sampling strategy that preserves structural diversity, reduces runtime by over 80× on OpenCitations Meta and GESIS, and is the only tractable method for obtaining complete metadata on ORKG.

By Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello
arXiv AI
Sep 3

Unifying biomedical knowledge in a modern multimodal graph

OptimusKG is a multimodal biomedical labeled property graph that integrates structured and semi‑structured resources to preserve detailed, type‑specific metadata across molecular, anatomical, clinical, and environmental domains. The graph contains nearly 191,000 nodes, over 21.8 million edges, and more than 67 million property instances derived from 18 ontologies, with a top‑level schema that enforces node and edge constraints while retaining granular provenance. Validation using the PaperQA3 agent found that 70.0% of sampled edges are supported by literature evidence, and the graph offers a standardized resource for machine learning, knowledge‑grounded retrieval, and hypothesis generation in biomedical research.

By Lucas Vittor, Ayush Noori, I\~naki Arango, Joaqu\'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik
arXiv AI
Aug 28

From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement

The paper introduces a novel LLM‑driven multi‑agent pipeline that converts relational databases into graph databases by standardizing table and column names and iteratively refining the graph schema through ETL, Analyzer, and Graph agents. The resulting graph database meets accuracy, groundedness, and faithfulness criteria and shows significant performance gains, achieving 85.6% Q&A accuracy—12.12% higher than an SQL agent on PostgreSQL—and reducing latency by roughly threefold on a BFSI dataset. This demonstrates an efficient, automated method for transforming tabular data into a more intuitive and faster‑executing graph format.

By Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy
arXiv AI
Jun 12

Agents-K1: Towards Agent-native Knowledge Orchestration

arXiv:2606. 13669v1 Announce Type: new Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration.

By Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi Xie, Xiang Zhuang, Yue Fan, Runmin Ma, Shiyang Feng, Xiangchao Yan, Anran Liu, Peng Ye, Wenlong Zhang, Shufei Zhang, Chunfeng Song, Fenghua Ling, Jie Zhou, Liang He, Bo Zhang, Lei Bai
arXiv AI
Sep 12

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.

By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv AI
Jul 28

Retrieval-Augmented Generation of Ontologies from Relational Databases

arXiv:2506. 01232v2 Announce Type: replace-cross Abstract: Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph population, ontology-based data access, graph-based learning, and automated reasoning.

By Nadeen Fathallah, Mojtaba Nayyeri, Athish A Yogi, Ratan Bahadur Thapa, Hans-Michael Tautenhahn, Anton Schnurpel, Steffen Staab
arXiv AI
Jun 8

MetaConfigurator: AI-Assisted RDF Authoring from JSON Data

arXiv:2606. 07094v1 Announce Type: cross Abstract: Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across systems due to lacking semantic interoperability.

By Felix Neubauer, Mahdi Jafarkhani, Kenichi Endo, J\"urgen Pleiss, Benjamin Uekermann