CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
CALM introduces a smooth trust gating mechanism for decentralized federated learning, replacing hard filtering of teacher models with class‑wise, sample‑wise, and label‑based weighting. It allows clients with heterogeneous architectures to distill knowledge from peers without a central server or shared data, even under severe non‑IID label skew. Experiments on CIFAR‑10, SVHN, OrganAMNIST, and Google Speech Commands show that CALM consistently outperforms uniform and hard‑filtered distillation and matches or exceeds other heterogeneous‑FL methods.
arXiv:2607. 01272v1 Announce Type: cross Abstract: Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and models must run on limited edge hardware.
arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...
arXiv:2609.36660v1 Announce Type: new Abstract: We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients witho...
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit subst...
The paper introduces a robust decentralized federated distillation approach that allows heterogeneous client models to collaborate using predictions on shared unlabeled public data. Each client evaluates received predictions across three modalities—class prediction, boundary decision, and prediction correlation—filters unreliable clients, assigns reliability-based weights, and constructs modality-specific teachers. The method validates distillation gradients against supervised gradients from private data, removes conflicting gradients, and proves convergence under Byzantine attacks, achieving improved accuracy on CIFAR-10 and CIFAR-100 under non‑IID data and malicious conditions.