arXiv Machine Learning

Incremental Evaluation and Training in Relational Deep Learning

arXiv:2608. 13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning.

arXiv AI
Jun 9

What Makes a Desired Graph for Relational Deep Learning?

arXiv:2606. 08491v1 Announce Type: new Abstract: Relational deep learning (RDL) converts relational databases (RDBs) into heterogeneous graphs, but graphs derived directly from database schemas are often not well suited for how graph neural networks (GNNs) perform relational reasoning.

By Yao Cheng, Siqiang Luo
arXiv AI
Sep 10

TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs

TTGBench is a new benchmark for temporal graph learning that evaluates both structural evolution and semantic drift in text‑attributed graphs. It includes six real‑world, text‑rich datasets with dual volatility and supports multi‑class and multi‑label temporal node classification, addressing gaps left by existing benchmarks. A comprehensive evaluation of 17 state‑of‑the‑art methods shows a clear divide: TGNNs excel at structural prediction but struggle with semantic tracking, while LLM‑based models perform better on semantic tasks but lag in structural prediction.

By Longfei Ma, Zemin Liu, Fei Wu
arXiv Machine Learning
Jun 5

The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning

arXiv:2606. 06397v1 Announce Type: new Abstract: Current evaluation practices in relational learning rely heavily on flat leaderboards that average performance across heterogeneous datasets, implicitly assuming a uniform underlying structure.

By Shuo Wang, Xiangyu Wang, Quanxin Wang, Bailin Wu, Bokui Wang, Shunyang Huang, Boyan Deng, Haonan Liu, Ruiyi Fang, Zhenxiang Xu, Boyu Wang, Zhao Kang
arXiv AI
Sep 15

Generalization Can Emerge in Tabular Foundation Models From a Single Table

The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.

By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas
arXiv Machine Learning
Aug 27

MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection

MetaSieve is a metapath selection layer that reduces subgraph size in relational deep learning by pruning uninformative metapaths using SQL join and aggregation statistics. It scores candidate metapath extensions with a lightweight function that favors informative yet lightweight paths, discarding those below a threshold. The method is independent of GNN parameters and, when applied to the RelBench benchmark, consistently cuts per‑epoch training time while preserving or improving accuracy.

By Fahim Shahriar Khan, Ashraf Aboulnaga
arXiv AI
Jun 30

Beyond IID: How General Are Tabular Foundation Models, Really?

arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.

By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter