The paper investigates whether training Mixture-of-Experts (MoE) routers can improve memory‑bandwidth locality on consumer GPUs. Using a new zero‑surgery telemetry tool, the authors measure that a large Qwen3‑235B model is bottlenecked by disk‑based expert access, and that an LRU cache can serve a majority of requests. They pre‑register experiments training 137 M‑parameter MoE models with locality‑aware losses, finding that while cache misses can drop up to 60 % (99 % static‑pin hit rate), every configuration fails to meet a strict 1 % perplexity threshold, indicating a tight coupling between cache efficiency and model quality.
By Shriniwas Ramesh Suram
arXiv:2608. 07157v1 Announce Type: new Abstract: Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone.
By Alireza Moayedikia, Alicia Troncoso Lora
arXiv:2607. 08784v1 Announce Type: cross Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge.
By Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong
Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measured decode is 0.
arXiv:2609.36433v1 Announce Type: new
Abstract: Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free gui...
By Yiming Liu, Ben Wan, Tongxuan Liu, Ao Wang, Yuqi Xiong, Fan Zhang, Hui Chen, Guiguang Ding
arXiv:2609.39074v1 Announce Type: cross
Abstract: Federated learning (FL) on memory-constrained edge devices faces a dilemma: first-order (FO) optimization (i.e., backpropagation) demands substantial...
By Qiyuan Chen, Xian Wu, Yanan Ma, Xianhao Chen