arXiv AI

On the Theoretical Limitations of Embedding-based Link Prediction

arXiv:2506. 22271v3 Announce Type: replace Abstract: Neural networks often map low-dimensional embeddings to high-dimensional output spaces.

arXiv Machine Learning
Aug 28

Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs

The paper introduces LitEm, a neural regression model that allows transductive knowledge graph embedding models to predict numerical attributes. LitEm achieves top or near‑top performance on most attributes across datasets such as FB15K‑237, YAGO15K, DB15K, and Mutagenesis. A co‑training framework further improves link prediction for bilinear models while enabling them to predict numerical attributes, demonstrating literal‑aware encoding of attribute information.

By Rupesh Sapkota, Louis Mozart Kamdem Teyou, Moshood Yekini, Caglar Demir, Axel-Cyrille Ngonga Ngomo
arXiv AI
4d ago

ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction

The paper introduces ImbalancE, an inference‑time latent search method that mitigates degree imbalance bias in Knowledge Graph Embedding models. It targets the problematic prediction of target entities with much lower degrees than anchor entities, a common issue in recommender systems and other applications. Experiments on benchmark datasets show that ImbalancE improves predictions on the most imbalanced triples compared to conventional methods.

By Alberto Bernardi, Luca Costabello, Christophe Gueret
arXiv Machine Learning
Jul 31

Fully Inductive Cardinality Estimation

arXiv:2607. 28311v1 Announce Type: cross Abstract: Query optimization of Basic Graph Patterns (BGP) SPARQL queries over Knowledge Graphs (KG) requires accurate cardinality estimation.

By Tim Schwabe, Lukas Ketzer, Maribel Acosta
arXiv Machine Learning
Aug 28

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

The paper presents a scalable graph neural network (GNN) system for friend recommendation on a massive social graph. It introduces two key design choices: multi-hash ID embeddings that shrink the embedding table by over 98% without hurting ranking quality, and a timestamp-sorted compressed sparse row (CSR) storage with binary search that reduces temporal neighbor sampling from linear to logarithmic time. Experiments on a 194‑million‑user, 28‑billion‑edge graph show that these techniques enable production‑grade performance, boosting friend additions by 16% and unique friend adders by 11.5% in an online A/B test.

By Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov
arXiv Machine Learning
Aug 20

GraphK: Variable-Size Graph Generation with Efficient Edge Construction

GraphK introduces an encoder‑sampler‑decoder framework that generates variable‑size graphs efficiently. It learns permutation‑invariant latent representations and samples new node embeddings via maximum likelihood, enabling both upscaling and downscaling of graph size. Edge construction uses KDTree‑based top‑k neighbor search in latent space, reducing computational cost while capturing graph properties.

By Resul Tugay, Eren Olu\u{g}, Elif Ak, Sule Gunduz Oguducu
arXiv Machine Learning
Jul 1

The Impact of Dimensionality on the Stability of Node Embeddings

arXiv:2604. 08492v2 Announce Type: replace Abstract: Previous work has shown that node embedding methods can produce different representations and downstream predictions across repeated training runs, even when trained on the same data with identical hyperparameters.

By Tobias Schumacher, Simon Reichelt, Markus Strohmaier
arXiv AI
Jun 17

Handling Feature Heterogeneity with Learnable Graph Patches

arXiv:2606. 17667v1 Announce Type: cross Abstract: In recent years, the rapid development of foundation models and graph pre-training technologies has spurred increasing interest in constructing a universal pre-trained graph model or Graph Foundation Model (GFM).

By Yifei Sun, Yang Yang, Xiao Feng, Zijun Wang, Haoyang Zhong, Chunping Wang, Lei Chen