arXiv Machine Learning

When Connected Does Not Mean Similar: Charting the Homophily Boundary of SNAP-KG for Streaming Entity Integration

SNAP‑KG is a framework that assigns new entities to semantic communities in a growing knowledge graph using only raw features, without graph access or retraining at inference time. The authors evaluate it on five multi‑view benchmarks and a large OGB‑WikiKG2 graph, all of which contain at least one homophilous view, and find strong performance. Extending the evaluation to three heterophilous graphs shows that when no homophilous view exists, clustering quality drops sharply for both SNAP‑KG and transductive baselines; the key factor is the homophily of the relation rather than the number of relations, and multi‑view fusion only helps if at least one homophilous relation is present. "whyItMatters":"The study reveals that SNAP‑KG’s effectiveness relies on the homophily assumption, highlighting a limitation for heterophilous knowledge graphs and suggesting a direction for future research on heterophily‑aware models."

arXiv Machine Learning
Aug 27

SNAP-KG: Streaming Node Assignment via Projection for Knowledge Graph Entity Integration

SNAP-KG is a framework for integrating newly arriving entities into knowledge graphs by assigning them to semantic communities using a projector that maps raw feature vectors into a learned embedding space. Unlike traditional multi-view graph clustering methods, SNAP-KG supports inductive inference for streaming entities without requiring graph access or model retraining. Experiments on five benchmark datasets and a large-scale KG show significant inference speedups and competitive clustering quality, while reducing candidate search for entity resolution and link prediction by up to 97%.

By Jui-Chien Lin, Mohammad Mohammadi Amiri, Oshani Seneviratne
arXiv AI
Sep 15

FedV-KGQA in Practice: Design Lessons and an Interactive Prototype

FedV-KGQA addresses multi‑hop question answering over vertically partitioned knowledge graphs where each silo holds disjoint relation types. The system trains local embeddings, concatenates silo‑specific entity views, anchors questions at a topic entity, and ranks candidates without sharing raw triples. Experiments show federated fusion nearly matches centralized accuracy, that anchoring and enrichment are more critical than embedding choice, and that the cheapest encoder depends on target accuracy.

By Md Saikat Islam Khan Bappy, Oshani Seneviratne
arXiv Machine Learning
Sep 3

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

The paper introduces GONE, a benchmark for evaluating knowledge unlearning in large language models using structured knowledge graphs, and presents Neighborhood-Expanded Distribution Shaping (NEDS), a framework that leverages graph connectivity to separate forgotten facts from their semantic neighborhood. GONE disentangles direct fact removal, reasoning-based leakage, and catastrophic forgetting, while NEDS achieves high unlearning efficacy and locality on LLaMA-3-8B and Mistral-7B. The dataset is publicly available on Hugging Face.

By Chahana Dahal, Ashutosh Balasubramaniam, Zuobin Xiong
arXiv AI
Aug 26

FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs

FedV-KGQA is a framework for multi-hop question answering over knowledge graphs that are vertically partitioned across different organizations. It allows entities to be shared while each silo retains disjoint sets of relations, using local graph enrichment and knowledge graph embeddings so that raw triples and relation parameters never leave the silo. The system includes a topic entity anchoring mechanism to ground questions in the correct graph neighborhood without runtime inter-silo communication, and it achieves performance close to centralized systems on three benchmarks, including 3-hop reasoning and robustness to embedding perturbations.

By Md Saikat Islam Khan Bappy, Oshani Seneviratne
arXiv AI
Jun 9

UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough

arXiv:2603. 29875v3 Announce Type: replace-cross Abstract: One of the key problems in Retrieval-augmented generation (RAG) systems is that chunk-based retrieval pipelines represent the source chunks as atomic objects, mixing the information contained within such a chunk into a single vector.

By Ryszard Tuora, Mateusz Gali\'nski, Micha{\l} Godziszewski, Micha{\l} Karpowicz, Mateusz Czy\.znikiewicz, Adam Kozakiewicz, Tomasz Zi\k{e}tkiewicz