arXiv Machine Learning

When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator Choice

arXiv:2606. 10249v1 Announce Type: new Abstract: We examine whether graph neural network (GNN) design rules generalize across benchmark families by studying aggregator selection (sum, mean, max) on 24 node-classification datasets spanning citation, heterophilic, LINKX Facebook-100, co-purchase, and co-authorship graphs.

arXiv Machine Learning
Sep 11

SynCo: Synthetic Community-Aware Attributed Graph Generator for Graph Neural Network Benchmarking

SynCo is a synthetic graph generator that lets users control node degree distributions and sub‑community structures, addressing limitations of existing generators that rely on power‑law distributions and lack flexibility. It is evaluated on graph mimicking, hyperparameter tuning, and node clustering, outperforming state‑of‑the‑art methods while preserving original data distributions. SynCo can generate large graphs with up to 2.1 million nodes.

By Guilherme Henrique Messias, Mariana Caravanti de Souza, Sylvia Iasulaitis, Alan Dem\'etrius Baria Valejo
arXiv Machine Learning
Aug 31

Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution

The paper introduces SEMGNN, an end‑to‑end self‑explainable multi‑label graph neural network that simultaneously classifies nodes and identifies edges contributing to each predicted label. Unlike post‑hoc explainers, SEMGNN jointly learns a predictor and a sparse edge‑mask explainer, leveraging label‑label correlations to improve classification and generate distinct, coherent explanations for each label. Experiments on synthetic and real‑world networks in social, entertainment, and life‑science domains demonstrate competitive predictive performance and more faithful, compact label‑conditioned explanations.

By Yingqi Feng, Yufei Tang, Min Shi, Xingquan Zhu
arXiv Machine Learning
Sep 11

HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation

HERALD is a new gradient‑free graph condensation framework that adapts node scoring and feature selection to a graph’s heterophily level. It selects features using a joint Fisher‑discriminability and activation‑density criterion, and scores nodes with a weighted combination of prototype representativeness, decision‑boundary proximity, and Local Intrinsic Dimensionality, where the weights depend on the heterophily ratio. The selected nodes are assembled into a condensed subgraph via score‑ordered BFS expansion, Personalized PageRank pruning, and class rebalancing, achieving comparable storage to BONSAI and outperforming state‑of‑the‑art condensers on heterophilic graphs while remaining competitive on homophilic ones across multiple GNN architectures.

By Sujan Chakraborty, Priyanka Saha, Saptarshi Bej
arXiv AI
Sep 1

HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs

HeTGB is a new benchmark for heterophilic text‑attributed graphs, consisting of five real‑world datasets where nodes have rich textual descriptions. It allows systematic evaluation of graph neural networks, pre‑trained language models, and co‑training methods on node classification. The benchmark highlights the utility of text attributes, the challenges of heterophilic TAGs, and the limitations of current models.

By Shujie Li, Yuxia Wu, Yuan Fang, Chuan Shi