The paper shows that a plain Transformer can serve as an effective graph learner by adding three lightweight modifications: simplified L₂ attention, adaptive RMS normalization, and an MLP-based positional encoding stem. These changes preserve token magnitude and enable the model to achieve high expressivity on graph benchmarks, outperforming more complex graph transformer variants. The results suggest that plain Transformers can act as a unified backbone for multimodal learning across language, vision, and graph domains.
By Liheng Ma, Soumyasundar Pal, Yingxue Zhang, Philip H. S. Torr, Mark Coates
arXiv:2602. 01553v3 Announce Type: replace-cross Abstract: Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies.
By Quang Truong, Yu Song, Donald Loveland, Mingxuan Ju, Tong Zhao, Neil Shah, Jiliang Tang
arXiv:2602.15239v3 Announce Type: replace
Abstract: Transformers have achieved remarkable success across domains, motivating the rise of Graph Transformers (GTs) as attention-based architectures for...
By Javier Porras-Valenzuela, Zhiyang Wang, Teresa Shang, Yusu Wang, Alejandro Ribeiro
The article surveys Dynamic Heterogeneous Graph Representation Learning (DHGRL), a field that tackles the challenges of modeling evolving, multi‑type networks. It introduces a unified definition covering both discrete‑time and continuous‑time DHGs, and proposes an algorithm‑centric taxonomy that groups methods into embedding‑based, GNN‑based, and Transformer‑based approaches, highlighting their biases toward temporal granularity. The survey also reviews key applications, datasets, benchmarks, and outlines future research directions.
By Huan Liu, Pengfei Jiao, Jie Yin, Hongjiang Chen, Zhidong Zhao
The paper investigates how the choice of graph tokenization affects transformer expressivity. It analyzes three tokenization families—spectral, random‑walk, and adjacency—showing that each induces different depth requirements and that some tokenizations are inherently lossy or ill‑conditioned for certain tasks. The authors prove lower bounds and impossibility results for converting between tokenizations and validate these findings with experiments on synthetic and real‑world data.
By Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna
arXiv:2609.07961v1 Announce Type: new
Abstract: Graph modeling, a crucial task for representing complex relationships in graph-structured data, has achieved significant success in recent years. Howev...
By Thanh-Dat Truong, Sarah Alharbi, Susan Gauch, Xinghui Zhao, Marios Savvides, Khoa Luu
arXiv:2412. 19419v2 Announce Type: replace-cross Abstract: Graph neural networks are deep neural networks designed for graphs with attributes attached to nodes or edges.
By James H. Tanis, Chris Giannella, Adrian V. Mariano, Daoud Meerzaman
arXiv:2603. 08825v2 Announce Type: replace-cross Abstract: Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on Graph Transformers or higher order architectures.
By Jay Revolinsky, Harry Shomer, Jiliang Tang
arXiv:2606. 05046v1 Announce Type: new Abstract: We introduce Graph Cascades, a mesoscopic rewiring strategy for Graph Neural Networks (GNNs) and Graph Transformers (GTs) that captures intermediate-scale graph structure beyond purely local edges or fully global attention.
By Meher Chaitanya, My Le, Luana Ruiz
HyPE-GT introduces a framework that generates learnable hyperbolic positional encodings for Graph Transformers, enabling the capture of complex hierarchical relationships in graph-structured data. Unlike traditional Euclidean encodings, HyPE’s hyperbolic encodings can be selected to suit specific downstream tasks and help mitigate oversmoothing in deep Graph Neural Networks. Experiments on molecular benchmarks and large-scale Open Graph Benchmark datasets demonstrate improved performance, while additional tests on Coauthor and Copurchase networks confirm HyPE’s effectiveness in controlling oversmoothing.
By Kushal Bose, Swagatam Das
arXiv:2605.06462v2 Announce Type: replace
Abstract: Progress in graph learning is hindered by benchmark practices that conflate the contributions of node features and graph structure, making it hard...
By Richard von Moos, Mathieu Alain, Bastian Rieck
The paper introduces higher-order positional encodings that enrich graph representations by incorporating topological information from lifted incidence structures, without altering existing graph learning backbones. It theoretically shows that these encodings can mix graph Laplacian frequencies beyond what scalar spectral filters achieve, and demonstrates their effectiveness on Graph Transformers for datasets like ZINC and synthetic benchmarks. The approach bridges graph positional encodings and topological deep learning, enabling standard models to exploit higher-order interactions.
By Caleb Stam, Aagrim Hoysal, Sanjukta Krishnagopal