arXiv Machine Learning

MTSSL: Meta-Thresholding Semi-Supervised Learning

arXiv:2607. 16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels.

arXiv AI
Aug 24

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

The paper introduces C-Score, a diagnostic framework for evaluating pseudo‑label‑based semi‑supervised learning (SSL) when unlabeled data may contain out‑of‑distribution (OOD) samples. C-Score assesses training behavior across prediction, feature representation, and optimization, using metrics such as PLE, CCI, Sem‑Drift, and Grad‑Align. Experiments on CIFAR‑10 and CIFAR‑100 with various OOD sources show that C‑Score detects hidden degradation that clean accuracy alone fails to reveal, highlighting the need for internal diagnostic signals in SSL robustness assessment.

By Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su
arXiv AI
Sep 24

AdaDim: Dimensionality Adaptation for SSL Representational Dynamics

AdaDim introduces a training strategy for self‑supervised learning that adaptively balances dimensionality increase and mutual information reduction. By gradually regularizing the projection head while encouraging feature decorrelation and sample uniformity, AdaDim achieves up to 3% performance gains over standard SSL baselines without relying on costly techniques such as queues or predictor networks. The method demonstrates that optimal SSL models do not simply maximize dimensionality or minimize mutual information, but find a trade‑off between the two.

By Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib
arXiv Machine Learning
Sep 23

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

The paper presents a practical approach to semi‑supervised federated learning for automatic speech recognition (ASR). It demonstrates that using a per‑client online teacher combined with a stabilizing server‑side anchor—where the server continues training on labeled data between rounds—significantly reduces divergence caused by pseudo‑label errors. The authors provide design guidelines that improve in‑domain performance by an average of 20.8 % and cross‑domain performance by 10.0 % over the best prior method, narrowing the gap to fully‑supervised federated learning.

By Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan "Honza" Silovsky, Tatiana Likhomanenko, Sheikh Shams Azam