arXiv Machine Learning

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

The paper presents a practical approach to semi‑supervised federated learning for automatic speech recognition (ASR). It demonstrates that using a per‑client online teacher combined with a stabilizing server‑side anchor—where the server continues training on labeled data between rounds—significantly reduces divergence caused by pseudo‑label errors. The authors provide design guidelines that improve in‑domain performance by an average of 20.8 % and cross‑domain performance by 10.0 % over the best prior method, narrowing the gap to fully‑supervised federated learning.

arXiv Machine Learning
Aug 14

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

arXiv:2608. 12773v1 Announce Type: cross Abstract: Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day.

By Ebenezer Tarubinga
arXiv Machine Learning
Sep 10

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

CALM introduces a smooth trust gating mechanism for decentralized federated learning, replacing hard filtering of teacher models with class‑wise, sample‑wise, and label‑based weighting. It allows clients with heterogeneous architectures to distill knowledge from peers without a central server or shared data, even under severe non‑IID label skew. Experiments on CIFAR‑10, SVHN, OrganAMNIST, and Google Speech Commands show that CALM consistently outperforms uniform and hard‑filtered distillation and matches or exceeds other heterogeneous‑FL methods.

By Yifan Ying, Qing Tian
arXiv Machine Learning
Jul 21

MTSSL: Meta-Thresholding Semi-Supervised Learning

arXiv:2607. 16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels.

By Shuyang Liu, Ziang Zeng, Ruiqiu Zheng, Jiazheng Wang, Zechen Liu, Wenxi Li, Zhou Yu
arXiv AI
Aug 25

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

FedCC is a new algorithm for distillation-based federated learning that tackles label distribution skew by allowing clients to mark ambiguous samples as 'unknown' instead of forcing a potentially wrong classification. By adding this extra class and calibrating pseudo-labels on a public dataset, FedCC balances confidence across majority and minority classes. Experiments show that FedCC outperforms existing methods, achieving 67.3% accuracy even when each client has data from only one of ten classes, whereas baselines drop to near-random performance.

By Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle
arXiv AI
Aug 18

Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank

arXiv:2608. 16681v1 Announce Type: cross Abstract: Although semi-supervised semantic segmentation ($\text{S}^4$) utilizes abundant unlabeled data to reduce manual labeling burdens, independent training of labeled and unlabeled data causes the former to dominate, which severely degrades pseudo-label quality.

By Shanwen Wang, Xin Sun, Danfeng Hong, Junyu Dong, Patrick Le Callet