arXiv:2605.06814v2 Announce Type: replace
Abstract: Graph neural networks (GNNs) increasingly rely on sophisticated architectures and training procedures to achieve desirable properties such as high...
By Debolina Halder Lina, Arlei Silva
arXiv:2604. 10882v2 Announce Type: replace-cross Abstract: Graph Neural Network pretraining is pivotal for leveraging unlabeled graph data.
By Yang Yan, Yunxuan Li, Qiuyan Wang, Tianjin Huang, Qiudong Yu
arXiv:2410. 00074v2 Announce Type: replace Abstract: A novel Learning-by-Education Node Community framework (LENC) for Collaborative Knowledge Distillation (CKD) is presented, which facilitates continual collective learning through effective knowledge exchanges among diverse deployed Deep Neural Network (DNN) peer nodes.
By Anestis Kaimakamidis, Ioannis Mademlis, Ioannis Pitas
arXiv:2609.13199v1 Announce Type: new
Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher mo...
By Dawen Jiang, Zhishu Shen, Zeyu Liu, Tiehua Zhang
arXiv:2607. 21885v1 Announce Type: new Abstract: Coarsening-based training for graph neural networks (GNNs), i.
By Guoming Li, Jian Yang, Xukun Wang, Zixiao Wang, Shangsong Liang, Yifan Chen
The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.
By Irene Trigueros-Lorca, Leonardo Concepci\'on, Christian Wagner, Isaac Triguero, Daniel Molina
arXiv:2608. 04377v1 Announce Type: cross Abstract: Hypergraph neural networks (HGNNs) have demonstrated remarkable capabilities in processing complex higher-order relationships.
By Mengyao Zhou, Zhiheng Zhou, Xiao Han, Guiying Yan
The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.
By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
arXiv:2601. 01484v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs.
By Itai Morad, Nir Shlezinger, Yonina C. Eldar
arXiv:2606.22975v2 Announce Type: replace
Abstract: Text-attributed graphs (TAGs) are widely used in many real-world domains, and learning on TAGs requires jointly modeling text semantics and graph s...
By Yeongho Kim, Yeonje Choi, Kijung Shin
arXiv:2609. 19210v1 Announce Type: cross Abstract: Graph neural networks are widely used for transductive node classification, with accuracy typically measured on randomly drawn train/validation/test splits.
By Naga Venkata Sai Jitin Jami, Thomas Altstidl, Sebastian Hoefler, Jonas Mueller, Dario Zanca, Bjoern Eskofier, Heike Leutheuser