The paper presents TASTE, a method that uses Bayesian optimization to tune batch size for on‑device edge learning, aiming to maximize hardware throughput while preserving accuracy. Experiments on devices like the Raspberry Pi 4 show that the tuned batch size, combined with gradient accumulation and linear learning‑rate scaling, can double training throughput compared to using the maximum batch size. In online continual learning, the optimal batch size also helps balance stability and plasticity, reducing catastrophic forgetting without sacrificing efficiency.
By Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. sLoTh freezes the backbone and limits plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, updating less than 1% of parameters without replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks while achieving roughly 6.5× lower energy consumption than dense vision transformers.
By Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur
arXiv:2606. 00888v1 Announce Type: cross Abstract: Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates.
By Qiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal, Maurice van Keulen, Elena Mocanu, Mykola Pechenizkiy, Decebal Constantin Mocanu, Torsten Hoefler
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. By freezing the backbone and limiting plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, sLoTh adapts to new tasks while updating less than 1% of the parameters and avoiding replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks and achieve roughly 6.5× lower energy consumption than dense vision transformers.
arXiv:2402.11215v4 Announce Type: replace
Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...
By Tim Tsz-Kit Lau, Han Liu, Mladen Kolar
arXiv:2505. 24852v3 Announce Type: replace-cross Abstract: On-device learning at the edge enables low-latency, private personalization with improved long-term robustness and reduced maintenance costs.
By Douwe den Blanken, Charlotte Frenkel
arXiv:2605. 11855v2 Announce Type: replace-cross Abstract: Sequence learning is dominated by Transformers and parallelizable recurrent neural networks (RNNs) such as state-space models, yet learning long-term dependencies remains challenging, and state-of-the-art designs trade power consumption for performance.
By Julien Brandoit, Arthur Fyon, Damien Ernst, Guillaume Drion
arXiv:2508. 03105v3 Announce Type: replace Abstract: We analyze the convergence behavior of stochastic gradient descent with momentum (SGDM) under dynamic learning-rate and batch-size schedules by introducing a novel and simpler Lyapunov function.
By Yuichi Kondo, Hideaki Iiduka
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
ExpTest is an autonomous learning‑rate controller that uses the training loss curve as an online signal to perform sequential statistical tests on theoretically motivated windows, detecting convergent behavior and triggering learning‑rate reductions. It combines a covariance‑based initial learning‑rate estimate, curvature‑motivated window sizing, and a two‑phase test‑driven decay, relying on the approximately exponential decay predicted under linearized network dynamics. Experiments on regression, classification, forecasting, and natural‑language tasks across various architectures show that ExpTest achieves competitive performance compared to hand‑tuned SGD baselines and recent learning‑rate‑free methods, without requiring manual initial learning‑rate selection or predefined scheduling.
By Zan Chaudhry, Naoko Mizuno
arXiv:2607. 15745v1 Announce Type: new Abstract: Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches.
By Anxhelo Shehu, Enes Stastoli, Arben Cela
arXiv:2609.17026v1 Announce Type: new
Abstract: Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a da...
By Yunxiang Fu, Meng Lou, Zicheng Liao, Yizhou Yu