arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
arXiv:2407. 05370v3 Announce Type: replace Abstract: Semi-supervised learning (SSL) algorithms often struggle to perform well when trained on imbalanced data.
By Zeju Li, Ying-Qiu Zheng, Chen Chen, Saad Jbabdi
arXiv:2601. 11670v3 Announce Type: replace-cross Abstract: Pseudo-label selection in semi-supervised learning is commonly driven by maximum-confidence thresholds, yet confidence alone can be unreliable under model overconfidence and class imbalance.
By Jinshi Liu, Lei He, Pan Liu
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2607. 16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels.
By Shuyang Liu, Ziang Zeng, Ruiqiu Zheng, Jiazheng Wang, Zechen Liu, Wenxi Li, Zhou Yu
arXiv:2603. 01346v2 Announce Type: replace Abstract: We revisit the framework of Smart PAC learning, which seeks supervised learners which compete with semi-supervised learners that are provided full knowledge of the marginal distribution on unlabeled data.
By Shaddin Dughmi, Alireza F. Pour
arXiv:2607. 13428v1 Announce Type: new Abstract: Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones.
By Xutao Wang, Hanting Chen, Tianyu Guo, Yunhe Wang
arXiv:2511. 22823v2 Announce Type: replace-cross Abstract: Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire.
By Miao Zhang, Junpeng Li, Changchun Hua, Yana Yang
arXiv:2607. 13550v1 Announce Type: cross Abstract: Boosting is one of the most successful learning techniques for standard classification and regression tasks.
By R\'emy Chapelle (CESP, CB, EVDG), Nicolas Vayatis (CB), Bruno Falissard (CESP), Mohammed Sedki (CESP)
arXiv:2606. 00512v1 Announce Type: new Abstract: In many modern machine learning pipelines, abundant pretrained representations serve as noisy proxy covariates, while task-specific labels remain scarce.
By Kwangho Kim, Jisu Kim
arXiv:2601. 17257v2 Announce Type: replace Abstract: We introduce a constrained optimization framework for training transformers that behave like optimization descent algorithms.
By Javier Porras-Valenzuela, Samar Hadou, Alejandro Ribeiro
arXiv:2607. 07671v1 Announce Type: new Abstract: Probabilistic circuits (PCs) can model complex joint distributions while supporting exact and efficient computation of many inference queries.
By Adrian Ciotinga, Yeming Dai, YooJung Choi