arXiv Machine Learning

Imbalanced Semi-Supervised Learning via Label Refinement and Threshold Adjustment

arXiv:2407. 05370v3 Announce Type: replace Abstract: Semi-supervised learning (SSL) algorithms often struggle to perform well when trained on imbalanced data.

arXiv Machine Learning
Jul 21

MTSSL: Meta-Thresholding Semi-Supervised Learning

arXiv:2607. 16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels.

By Shuyang Liu, Ziang Zeng, Ruiqiu Zheng, Jiazheng Wang, Zechen Liu, Wenxi Li, Zhou Yu
arXiv AI
Aug 24

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

The paper introduces C-Score, a diagnostic framework for evaluating pseudo‑label‑based semi‑supervised learning (SSL) when unlabeled data may contain out‑of‑distribution (OOD) samples. C-Score assesses training behavior across prediction, feature representation, and optimization, using metrics such as PLE, CCI, Sem‑Drift, and Grad‑Align. Experiments on CIFAR‑10 and CIFAR‑100 with various OOD sources show that C‑Score detects hidden degradation that clean accuracy alone fails to reveal, highlighting the need for internal diagnostic signals in SSL robustness assessment.

By Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su
arXiv AI
Aug 25

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

FedCC is a new algorithm for distillation-based federated learning that tackles label distribution skew by allowing clients to mark ambiguous samples as 'unknown' instead of forcing a potentially wrong classification. By adding this extra class and calibrating pseudo-labels on a public dataset, FedCC balances confidence across majority and minority classes. Experiments show that FedCC outperforms existing methods, achieving 67.3% accuracy even when each client has data from only one of ten classes, whereas baselines drop to near-random performance.

By Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle
arXiv Machine Learning
Sep 25

Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution

The study evaluates pseudo‑labeling for semi‑supervised learning on Android malware attribution using six classifiers. Results show that the benefit of SSL varies strongly by classifier: SVM gains the most, LightGBM improves modestly, and Random Forest can be harmed at low label ratios. The approach particularly helps hard‑to‑classify families and achieves near‑optimal performance with about 800 labeled samples.

By Md Rafid Islam, Zahid Hasan, Hafiz Abdur Rahman