arXiv Machine Learning

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

CALM introduces a smooth trust gating mechanism for decentralized federated learning, replacing hard filtering of teacher models with class‑wise, sample‑wise, and label‑based weighting. It allows clients with heterogeneous architectures to distill knowledge from peers without a central server or shared data, even under severe non‑IID label skew. Experiments on CIFAR‑10, SVHN, OrganAMNIST, and Google Speech Commands show that CALM consistently outperforms uniform and hard‑filtered distillation and matches or exceeds other heterogeneous‑FL methods.

arXiv Machine Learning
Sep 22

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

arXiv:2609.23697v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...

By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv Machine Learning
Sep 10

Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration

The paper introduces a robust decentralized federated distillation approach that allows heterogeneous client models to collaborate using predictions on shared unlabeled public data. Each client evaluates received predictions across three modalities—class prediction, boundary decision, and prediction correlation—filters unreliable clients, assigns reliability-based weights, and constructs modality-specific teachers. The method validates distillation gradients against supervised gradients from private data, removes conflicting gradients, and proves convergence under Byzantine attacks, achieving improved accuracy on CIFAR-10 and CIFAR-100 under non‑IID data and malicious conditions.

By Xiao Ma, Hong Shen, Hui Tian, Wei Ke, Wenqi Lyu
arXiv Machine Learning
Sep 23

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

The paper presents a practical approach to semi‑supervised federated learning for automatic speech recognition (ASR). It demonstrates that using a per‑client online teacher combined with a stabilizing server‑side anchor—where the server continues training on labeled data between rounds—significantly reduces divergence caused by pseudo‑label errors. The authors provide design guidelines that improve in‑domain performance by an average of 20.8 % and cross‑domain performance by 10.0 % over the best prior method, narrowing the gap to fully‑supervised federated learning.

By Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan "Honza" Silovsky, Tatiana Likhomanenko, Sheikh Shams Azam
arXiv AI
Sep 4

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

The paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level before applying dense supervision in on-policy distillation. TGOPD uses verifier-scored teacher probes to decide whether to route a prompt to dense OPD or to a verifier-grounded alternative. Experiments on 4B and 35B models across mathematics, code, and instruction tasks show TGOPD outperforms vanilla OPD and improves teacher GPU utilization from 9.8% to 78.9% in a 4B single-domain run.

By Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan
arXiv Machine Learning
4d ago

Byzantine-Robust Federated Representation Learning

arXiv:2609.36660v1 Announce Type: new Abstract: We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients witho...

By Leonardo F. Toso, James Anderson, Rafael Pinot, Nirupam Gupta