arXiv:2609.14652v1 Announce Type: cross
Abstract: Large Language Model (LLM) applications often transfer domain concepts into the model's context informally, through prompt prose, schema dumps, and e...
By Blake G. Fitch
The paper introduces EXYGEN, a framework that enables conversational access to large knowledge graphs by combining VoID descriptions, ShEx schemas, retrieved triples, and example question‑query pairs in a retrieval‑augmented generation pipeline. On the SciQA benchmark, this approach achieves an exact‑match score of 0.419 without fine‑tuning any large language model, and shows that larger general‑purpose LLMs can outperform smaller code‑specialized ones when provided sufficient context. To scale metadata generation for very large KGs, the authors propose a predicate‑coverage‑aware parallel graph sampling strategy that preserves structural diversity, reduces runtime by over 80× on OpenCitations Meta and GESIS, and is the only tractable method for obtaining complete metadata on ORKG.
By Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello
SPARQL-LLM is an open‑source, triplestore‑agnostic system that generates SPARQL queries from natural language using lightweight metadata and dedicated components for indexing, prompt building, and execution. It achieves up to 59 % higher F1 scores than the next best system on a multilingual challenge and on bioinformatics knowledge graphs, while being up to 27 × faster and costing no more than $0.01 per question. The project is publicly available on GitHub and is already deployed on real‑world decentralized knowledge graphs such as expasy.org/chat.
By Panayiotis Smeros, Vincent Emonet, Ruijie Wang, Ana-Claudia Sima, Tarcisio Mendes de Farias
arXiv:2507. 21438v2 Announce Type: replace Abstract: Ontologies and knowledge graphs require continuous evolution to remain comprehensive and accurate, but manual curation is labor intensive.
By Vishal Raman, Vijai Aravindh R, Abhijith Ragav
Build2SPARQL is a large-scale benchmark dataset for translating natural-language questions into SPARQL queries over building knowledge graphs. The dataset is generated by a KG‑grounded pipeline that produces 6,136 executable SPARQL queries and 30,680 corresponding natural-language questions across six query-pattern families and five vocabulary registers, covering 201 building KGs. Human validation shows high semantic fidelity, naturalness, and operational plausibility, and retrieval‑augmented evaluation demonstrates significant accuracy gains for open‑weight language models.
By Wooyoung Jung
arXiv:2608. 07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and semantically faithful.
By Tommaso Soru, Abdulsobur Oyewale