arXiv Machine Learning

CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

arXiv Machine Learning
Sep 10

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

CALM introduces a smooth trust gating mechanism for decentralized federated learning, replacing hard filtering of teacher models with class‑wise, sample‑wise, and label‑based weighting. It allows clients with heterogeneous architectures to distill knowledge from peers without a central server or shared data, even under severe non‑IID label skew. Experiments on CIFAR‑10, SVHN, OrganAMNIST, and Google Speech Commands show that CALM consistently outperforms uniform and hard‑filtered distillation and matches or exceeds other heterogeneous‑FL methods.

By Yifan Ying, Qing Tian
arXiv Machine Learning
Sep 22

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...

By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv Machine Learning
4d ago

Byzantine-Robust Federated Representation Learning

arXiv:2609.36660v1 Announce Type: new Abstract: We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients witho...

By Leonardo F. Toso, James Anderson, Rafael Pinot, Nirupam Gupta
arXiv Machine Learning
Sep 10

Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration

The paper introduces a robust decentralized federated distillation approach that allows heterogeneous client models to collaborate using predictions on shared unlabeled public data. Each client evaluates received predictions across three modalities—class prediction, boundary decision, and prediction correlation—filters unreliable clients, assigns reliability-based weights, and constructs modality-specific teachers. The method validates distillation gradients against supervised gradients from private data, removes conflicting gradients, and proves convergence under Byzantine attacks, achieving improved accuracy on CIFAR-10 and CIFAR-100 under non‑IID data and malicious conditions.

By Xiao Ma, Hong Shen, Hui Tian, Wei Ke, Wenqi Lyu
arXiv Machine Learning
Aug 31

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.

By Yiming Xie, Lili Su, Ningfang Mi