The paper presents TASTE, a method that uses Bayesian optimization to tune batch size for on‑device edge learning, aiming to maximize hardware throughput while preserving accuracy. Experiments on devices like the Raspberry Pi 4 show that the tuned batch size, combined with gradient accumulation and linear learning‑rate scaling, can double training throughput compared to using the maximum batch size. In online continual learning, the optimal batch size also helps balance stability and plasticity, reducing catastrophic forgetting without sacrificing efficiency.
By Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. sLoTh freezes the backbone and limits plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, updating less than 1% of parameters without replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks while achieving roughly 6.5× lower energy consumption than dense vision transformers.
By Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur
arXiv:2606. 00888v1 Announce Type: cross Abstract: Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates.
By Qiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal, Maurice van Keulen, Elena Mocanu, Mykola Pechenizkiy, Decebal Constantin Mocanu, Torsten Hoefler
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. By freezing the backbone and limiting plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, sLoTh adapts to new tasks while updating less than 1% of the parameters and avoiding replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks and achieve roughly 6.5× lower energy consumption than dense vision transformers.
arXiv:2402.11215v4 Announce Type: replace
Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...
By Tim Tsz-Kit Lau, Han Liu, Mladen Kolar
arXiv:2505. 24852v3 Announce Type: replace-cross Abstract: On-device learning at the edge enables low-latency, private personalization with improved long-term robustness and reduced maintenance costs.
By Douwe den Blanken, Charlotte Frenkel