arXiv Machine Learning

Self-Soupervision: Cooking Model Soups without Labels

arXiv:2602. 02890v2 Announce Type: replace Abstract: Model soups are strange and strangely effective combinations of parameters.

arXiv AI
Aug 24

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

The paper introduces C-Score, a diagnostic framework for evaluating pseudo‑label‑based semi‑supervised learning (SSL) when unlabeled data may contain out‑of‑distribution (OOD) samples. C-Score assesses training behavior across prediction, feature representation, and optimization, using metrics such as PLE, CCI, Sem‑Drift, and Grad‑Align. Experiments on CIFAR‑10 and CIFAR‑100 with various OOD sources show that C‑Score detects hidden degradation that clean accuracy alone fails to reveal, highlighting the need for internal diagnostic signals in SSL robustness assessment.

By Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su
arXiv AI
Aug 25

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

FedCC is a new algorithm for distillation-based federated learning that tackles label distribution skew by allowing clients to mark ambiguous samples as 'unknown' instead of forcing a potentially wrong classification. By adding this extra class and calibrating pseudo-labels on a public dataset, FedCC balances confidence across majority and minority classes. Experiments show that FedCC outperforms existing methods, achieving 67.3% accuracy even when each client has data from only one of ten classes, whereas baselines drop to near-random performance.

By Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle
arXiv AI
Sep 18

Pre-train to Gain: Robust Learning Without Clean Labels

The paper proposes a method that first pre‑trains a feature extractor on the target dataset using in‑domain self‑supervised learning (SSL) without labels, then performs standard supervised training on the same noisy dataset. This two‑stage approach eliminates the need for a clean label subset and consistently improves classification accuracy and label‑error detection across synthetic and real‑world noise, especially as noise rates increase. Experiments show that the method matches or surpasses ImageNet and DinoV2 pre‑training, particularly under high noise conditions.

By David Szczecina, Nicholas Pellegrino, Paul Fieguth
arXiv Machine Learning
Jul 21

MTSSL: Meta-Thresholding Semi-Supervised Learning

arXiv:2607. 16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels.

By Shuyang Liu, Ziang Zeng, Ruiqiu Zheng, Jiazheng Wang, Zechen Liu, Wenxi Li, Zhou Yu