arXiv:2607. 07156v1 Announce Type: new Abstract: Different optimizers have different update biases, but these biases are usually implicit.
By Zhang Gongyue, Liu Donghan, Ren Weihong, Sheng Yixuan, Wang Zhiyong, Liu Honghai
arXiv:2608. 07157v1 Announce Type: new Abstract: Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone.
By Alireza Moayedikia, Alicia Troncoso Lora
arXiv:2608. 11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting.
By Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
arXiv:2512. 12816v2 Announce Type: replace Abstract: We study how to allocate resources for training and deployment of machine learning (ML) models under concept drift and limited budgets.
By Hasan Burhan Beytur, Haris Vikalo, Kevin S Chan, Gustavo de Veciana
arXiv:2608. 01032v1 Announce Type: new Abstract: Training error is what we can observe on a training set; test error is the quantity we actually care about.
By Gireeja Ranade, Anant Sahai
arXiv:2606. 00340v1 Announce Type: new Abstract: We study optimal learning-rate selection in two-layer and three-layer linear neural networks trained to learn linear target functions.
By Tianyu Pang, Vignesh Kothapalli, Shenyang Deng, Haohui Wang, Dawei Zhou, Yaoqing Yang