The paper introduces EXYGEN, a framework that enables conversational access to large knowledge graphs by combining VoID descriptions, ShEx schemas, retrieved triples, and example question‑query pairs in a retrieval‑augmented generation pipeline. On the SciQA benchmark, this approach achieves an exact‑match score of 0.419 without fine‑tuning any large language model, and shows that larger general‑purpose LLMs can outperform smaller code‑specialized ones when provided sufficient context. To scale metadata generation for very large KGs, the authors propose a predicate‑coverage‑aware parallel graph sampling strategy that preserves structural diversity, reduces runtime by over 80× on OpenCitations Meta and GESIS, and is the only tractable method for obtaining complete metadata on ORKG.
By Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello
arXiv:2609.01525v1 Announce Type: cross
Abstract: A durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data....
By Gene Zhang
text2ql is an open‑source Python framework that enables natural language querying of databases without relying on large language models at query time. It uses a language‑agnostic intermediate representation (QueryIR) and a pluggable renderer to support both SQL and GraphQL targets through a single seven‑stage detection pipeline. In deterministic mode, it achieves 100% execution accuracy with a median latency of 3.2 ms, while the LLM‑backed mode delivers 62‑70% exact match and 84‑91% execution accuracy on benchmark samples.
By Ritesh Kumar
Build2SPARQL is a large-scale benchmark dataset for translating natural-language questions into SPARQL queries over building knowledge graphs. The dataset is generated by a KG‑grounded pipeline that produces 6,136 executable SPARQL queries and 30,680 corresponding natural-language questions across six query-pattern families and five vocabulary registers, covering 201 building KGs. Human validation shows high semantic fidelity, naturalness, and operational plausibility, and retrieval‑augmented evaluation demonstrates significant accuracy gains for open‑weight language models.
By Wooyoung Jung
arXiv:2508. 01815v2 Announce Type: replace-cross Abstract: Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs.
By Yang Zhao, Chengxiao Dai, Yue Xiu, Dusit Niyato
arXiv:2502. 11201v3 Announce Type: replace-cross Abstract: NoSQL databases are core data infrastructure, yet natural-language access to them remains underdeveloped: correct query generation must recover how a non-relational data model represents entities, nested paths, arrays, missing fields, and dynamic keys.
By Jinwei Lu, Jiawei Lu, Chen Zhang, Zhiqian Qin, Haodi Zhang, Yuanfeng Song, Raymond Chi-Wing Wong