arXiv AI

FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity

arXiv:2607. 03763v1 Announce Type: cross Abstract: Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD-based training.

arXiv AI
2d ago

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.

By Junkang Liu
arXiv AI
Aug 19

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.

By Akiyoshi Tomihari, Issei Sato
arXiv Machine Learning
1d ago

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.

By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li
arXiv Machine Learning
Aug 28

Federated Adversarial Training with Transformers

The paper investigates federated adversarial training (AT) for vision transformers, a topic not previously explored in federated learning (FL). It evaluates various transformer architectures and aggregation strategies, and introduces FedWAvg, an extension of FedAvg that weights client updates based on similarity of their last-layer representations. Experiments demonstrate that FedWAvg yields higher robust accuracy than existing aggregation methods in non‑IID settings.

By Ahmed Aldahdooh, Wassim Hamidouche, Olivier D\'eforges