Towards Data Science

Recursive CTEs: SQL’s Hidden Graph Traversal Engine

The article titled "Recursive CTEs: SQL’s Hidden Graph Traversal Engine" offers a practical guide for working with hierarchies in SQL. It explains how to navigate these structures, find routes, detect cycles, and calculate degrees of separation using recursive common table expressions. The guide is aimed at readers looking to leverage SQL’s built‑in graph traversal capabilities.

arXiv AI
Aug 28

From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement

The paper introduces a novel LLM‑driven multi‑agent pipeline that converts relational databases into graph databases by standardizing table and column names and iteratively refining the graph schema through ETL, Analyzer, and Graph agents. The resulting graph database meets accuracy, groundedness, and faithfulness criteria and shows significant performance gains, achieving 85.6% Q&A accuracy—12.12% higher than an SQL agent on PostgreSQL—and reducing latency by roughly threefold on a BFSI dataset. This demonstrates an efficient, automated method for transforming tabular data into a more intuitive and faster‑executing graph format.

By Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy
arXiv AI
Jul 28

Answering Path Queries under Linear and Guarded Existential Rules

arXiv:2607. 22636v1 Announce Type: new Abstract: Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology.

By Jean-Fran\c{c}ois Baget (LIRMM, Inria, University of Montpellier, CNRS, France), Meghyn Bienvenu (Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, France), Marie-Laure Mugnier (LIRMM, Inria, University of Montpellier, CNRS, France), Micha\"el Thomazo (Inria, DIENS, ENS, PSL University, CNRS, France)
arXiv AI
Sep 23

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

The paper evaluates seven graph database engines, including Corvic AI, on a synthetic biomedical property graph with 1.02 million nodes and 5.34 million rows. It benchmarks query latency, bulk‑ingest throughput, point‑update latency, and correctness across a twenty‑query workload that covers neighborhood lookups, bounded paths, set intersections, anti‑joins, aggregation, ranking, temporal filters, full scans, and relational joins. The study finds that no single engine is universally fastest; performance depends on query shape, and the largest cost difference arises from bulk‑ingest throughput, which varies by three orders of magnitude and dominates total cost for workloads with fewer than about 10⁵ queries per data refresh.

By Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach
Towards Data Science
Aug 20

Making the Knowledge Layer a Graph You Actually Traverse

The article discusses why retrieval quality should be inherent to the system rather than dependent on how a question is phrased. It proposes reconstructing the knowledge layer by performing graph traversal on every query, incorporating bitemporal edges, and applying a two‑threshold entity resolution approach. These techniques aim to make the knowledge graph more dynamic and responsive to user queries.

By Miodrag Cekikj
Towards Data Science
Sep 27

GraphRAG with TypeSafe Jev: A System One Approach to Scalable Knowledge Graphs

The article discusses how calibrated decision models can manage high‑frequency graph decisions while large language models (LLMs) concentrate on reasoning, synthesis, and open‑ended generation. It introduces GraphRAG with TypeSafe Jev as a system‑one approach to building scalable knowledge graphs. The focus is on separating decision‑making from generative tasks to improve efficiency and reliability.

By Partha Sarkar