The paper presents TASTE, a method that uses Bayesian optimization to tune batch size for on‑device edge learning, aiming to maximize hardware throughput while preserving accuracy. Experiments on devices like the Raspberry Pi 4 show that the tuned batch size, combined with gradient accumulation and linear learning‑rate scaling, can double training throughput compared to using the maximum batch size. In online continual learning, the optimal batch size also helps balance stability and plasticity, reducing catastrophic forgetting without sacrificing efficiency.
By Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann
arXiv:2509. 13211v4 Announce Type: replace Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models.
By Irene Testa, Luigi Quarantiello, Eric Nuertey Coleman, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. By freezing the backbone and limiting plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, sLoTh adapts to new tasks while updating less than 1% of the parameters and avoiding replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks and achieve roughly 6.5× lower energy consumption than dense vision transformers.
arXiv:2603. 10046v2 Announce Type: replace Abstract: Wearable sensors in Internet of Things (IoT) ecosystems increasingly support applications such as remote health monitoring, elderly care, and smart home automation, all of which rely on robust human activity recognition (HAR).
By Reza Rahimi Azghan, Gautham Krishna Gudur, Mohit Malu, Edison Thomaz, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh
The paper introduces sLoTh, a parameter‑efficient continual learning framework for sparse event‑based vision transformers. sLoTh freezes the backbone and limits plasticity to low‑rank attention updates (seLoRA) and shared neuronal threshold modulation, updating less than 1% of parameters without replay buffers. Experiments on CIFAR‑100, Tiny‑ImageNet, ImageNet‑100, and ImageNet‑R show competitive rehearsal‑free performance across up to 100 tasks while achieving roughly 6.5× lower energy consumption than dense vision transformers.
By Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur
arXiv:2604. 07396v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference on edge Neural Processing Units (NPUs) is fundamentally constrained by limited on-chip memory capacity.
By Jintao Zhang, Xuanyao Fong
The paper presents a novel multi‑exit computational scheme for TinyML on an ultra‑low‑power GAP9 SoC, adding confidence‑based gating points to a MobileNetV2 CNN for ImageNet‑100. By allowing inference to stop early, the approach cuts average MAC operations by 41 % (from 313 MMAC to 185 MMAC), reduces inference time by 29 % (49 ms to 35 ms), and saves 24 % in energy (2.1 mJ to 1.6 mJ per frame) with only a ~1 % drop in accuracy. Compared to a state‑of‑the‑art adaptive CNN on the same hardware, the method more than doubles computational efficiency, raising MAC/cycle from 8.1 to 17.2.
By Luca Crupi, Lorenzo Lamberti, Alessandro Giusti, Daniele Palossi
arXiv:2607. 29353v1 Announce Type: cross Abstract: With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.
By Douwe den Blanken, Martin Lefebvre, Charlotte Frenkel
The paper presents a generative continual learning framework that extends self‑organizing maps (SOMs) with learned distributional statistics and encoder–decoder models for class‑incremental learning. By storing running means, variances, and covariances for each SOM unit, the method can generate synthetic samples for replay without storing raw data, enabling exemplar‑free learning. Experiments on CIFAR‑10, CIFAR‑100, and TinyImageNet show competitive or superior performance compared to state‑of‑the‑art memory‑based and memory‑free methods, and the approach also allows easy visualization and post‑training generative use.
By Pujan Thapa, Alexander Ororbia, Travis Desell
arXiv:2606. 09960v1 Announce Type: cross Abstract: We present HydraCIL, a decoupled continual learning model based on prototype-guided multi-head classifiers, targeting sustainable deployment in embedded and resource-constrained environments.
By Daniel Vila-Cruz, Laura Mor\'an-Fern\'andez, Ver\'onica Bol\'on-Canedo
arXiv:2603. 01761v2 Announce Type: replace-cross Abstract: Foundation models have transformed machine learning through large-scale pretraining and increased test-time compute.
By Vaggelis Dorovatas, Malte Schwerin, Andrew D. Bagdanov, Lucas Caccia, Antonio Carta, Laurent Charlin, Barbara Hammer, Tyler L. Hayes, Timm Hess, Christopher Kanan, Dhireesha Kudithipudi, Xialei Liu, Vincenzo Lomonaco, Jorge Mendez-Mendez, Darshan Patil, Ameya Prabhu, Elisa Ricci, Tinne Tuytelaars, Gido M. van de Ven, Liyuan Wang, Joost van de Weijer, Jonghyun Choi, Martin Mundt, Rahaf Aljundi
arXiv:2505. 24852v3 Announce Type: replace-cross Abstract: On-device learning at the edge enables low-latency, private personalization with improved long-term robustness and reduced maintenance costs.
By Douwe den Blanken, Charlotte Frenkel