arXiv AI

WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement

WiCleanData is a refined version of Wikidata that eliminates type inconsistencies and constraint violations. The authors built an automated pipeline that cleans the taxonomy with language‑model assistance, aggregates type constraints hierarchically, and filters facts to ensure no type violations remain. The resulting knowledge graph is publicly available through a web interface for easy exploration and downstream use.

arXiv Computation and Language
Sep 25

CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding

The paper introduces CodeGraph, an open‑taxonomy knowledge graph that semantically annotates source code by extracting entities such as algorithms, paradigms, design patterns, and application domains from millions of files. Using a specialized large language model and a three‑stage Wikidata linking process, the authors ground these entities in Wikidata and construct a graph with about 158 million nodes and 1 billion typed edges across 14 programming languages. A quality‑assurance protocol combining human evaluation and an LLM‑as‑a‑judge filter quantifies annotation precision.

By Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
arXiv AI
Sep 21

From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata

The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.

By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
arXiv AI
Aug 26

Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA

The paper introduces Constrained Entity Selection under Partial Knowledge (CES-PK), a framework for improving large language model (LLM) based knowledge graph question answering (KGQA) by filtering candidate answers with lightweight symbolic constraints instead of full semantic parsing. CES-PK uses a three-valued constraint semantics—satisfied, violated, unknown—to handle incomplete knowledge graphs and avoid incorrect rejections under open‑world assumptions. Experiments on the Hetionet biomedical knowledge graph show that applying type, relation, and exclusion constraints increases precision while preserving recall, and that satisfied constraints can be used to rank remaining candidates.

By Emanuel Kitzelmann