ClusterAttention is a training‑free technique that speeds up bidirectional attention by recursively clustering keys and queries into fixed‑size, power‑of‑two blocks, enabling block‑sparse attention to match dense attention latency on GPUs. The method derives error bounds for sparse attention, showing tighter clusters can reduce error when compensated via centroids, and demonstrates significant speedups—up to six‑fold on large tabular data and 1.8× on video generation—while preserving over 99% of dense accuracy.
By Kasper Nordenram, Amelie Dittmann
arXiv:2603. 11475v2 Announce Type: replace Abstract: Accurate prediction of multivariate time series is essential for emerging network intelligent control, observability, and management functions.
By Yufeng Xin, Ethan Fan
arXiv:2606. 01660v1 Announce Type: new Abstract: Pre-propagation graph neural networks (PPGNNs) push all graph-dependent computation into a preprocessing step and train only on the resulting dense hop features, which makes them highly scalable.
By Zichao Yue, Zhiru Zhang
arXiv:2505. 13102v4 Announce Type: replace-cross Abstract: Unlike conventional "black-box" transformers with classical self-attention mechanism, we build a lightweight and interpretable transformer-like neural net by unrolling a mixed-graph-based optimization algorithm to forecast traffic with spatial and temporal dimensions.
By Ji Qi, Tam Thuc Do, Mingxiao Liu, Zhuoshi Pan, Yuzhe Li, Gene Cheung, H. Vicky Zhao
arXiv:2608. 14177v1 Announce Type: cross Abstract: Deep spatiotemporal models integrating graph convolutions and attention mechanisms have demonstrated excellent performance in network-level traffic flow prediction, owing to their exceptional ability to capture complex spatiotemporal dependencies.
By Xuanmian He, Can Li, Wanjing Ma
STHMoE is a Spatio‑Temporal Hypergraph‑Enhanced Mixture of Experts framework designed for urban traffic forecasting. It separates traffic dynamics into frequency‑, time‑, spatial‑, and higher‑order representations, each handled by a prompt‑guided expert built on a partially frozen large language model. The higher‑order expert uses an adaptive hypergraph module to learn evolving spatial structures, while an entropy‑aware router balances expert usage and fuses outputs, achieving competitive results on ten real‑world traffic benchmarks.
By Jiawen Chen, Qi Shao, Yongjian Chang, Mingtong Zhou, Duxin Chen, Wenwu Yu