arXiv:2609.15975v1 Announce Type: cross
Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
By Shwai He, Haichao Zhang, Shen Yan
arXiv:2509. 22082v3 Announce Type: replace Abstract: Federated learning enables distributed information sharing and collaborative model training without exposing raw client data.
By Li Xia, Jing Yu, Zheng Liu, Sili Huang, Wei Tang, Xuan Liu
arXiv:2606. 13896v1 Announce Type: cross Abstract: Self-supervised geospatial foundation models (GeoFMs) learn transferable representations from remote sensing data, but their downstream behavior is difficult to characterize.
By Julia Romero, Qin Lv, Morteza Karimzadeh
FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.
By Junkang Liu
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.
By Akiyoshi Tomihari, Issei Sato
FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.
By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li
arXiv:2607. 02612v1 Announce Type: cross Abstract: Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when many regions are redundant or uninformative.
By Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz
arXiv:2607. 28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets.
By Claudia Grosser, Maike Heuer, Denis Krompass, Thomas A. Runkler
arXiv:2609.37033v1 Announce Type: new
Abstract: Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but sta...
By Mengjun Yi, Huaian Gu, Yinghao Ai, Furao Shen, Jian Zhao
arXiv:2606. 04074v1 Announce Type: cross Abstract: Adaptive patching is a recent and compelling proposal for time-series Transformers: allocate finer patches where the sequence looks locally informative.
By Federico Zucchi, Yi Xie, Chao Zhang, Keyuan Luo, Thomas Lampert, Ziyue Li
The paper investigates federated adversarial training (AT) for vision transformers, a topic not previously explored in federated learning (FL). It evaluates various transformer architectures and aggregation strategies, and introduces FedWAvg, an extension of FedAvg that weights client updates based on similarity of their last-layer representations. Experiments demonstrate that FedWAvg yields higher robust accuracy than existing aggregation methods in non‑IID settings.
By Ahmed Aldahdooh, Wassim Hamidouche, Olivier D\'eforges
arXiv:2607.07494v2 Announce Type: replace-cross
Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision fo...
By Jieying Wang, Zizhong Wang, Fangru Linghu, Shuyuan Fan, Jiajia Li, Zhao Zhang