arXiv:2608. 09417v2 Announce Type: replace Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
By Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang
arXiv:2608. 09417v1 Announce Type: new Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.
By Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang
arXiv:2606. 11319v1 Announce Type: new Abstract: Learning from imperfect data is a central theme in machine learning, connecting practical questions of robustness to fundamental questions of learnability.
By Justin Tahmassebpur, Asadullah Bhuiyan, Hyejin Kim, Omri Lesser
arXiv:2607. 12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models.
By Ayush Karmacharya (Purdue University), Luke Luschwitz (Purdue University), Lucia Romero (Purdue University), Yanan Niu (EPFL), Joseph Campbell (Purdue University)
arXiv:2508. 09697v4 Announce Type: replace Abstract: Noisy labels are inevitable in real-world multimedia applications.
By Xinlei Zhang, Fan Liu, Chuanyi Zhang, Xiaoying Ji, Wenhui Wang, Wei Zhou, Yuhui Zheng
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat