Focused PU learning from imbalanced data
arXiv:2605.14467v2 Announce Type: replace Abstract: We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as dise...
arXiv:2607. 13428v1 Announce Type: new Abstract: Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones.
arXiv:2605.14467v2 Announce Type: replace Abstract: We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as dise...
arXiv:2606. 03332v1 Announce Type: new Abstract: Probabilistic models are typically trained using task-agnostic objectives like log-loss, which can lead to significant errors in downstream estimation.
The paper introduces a method for maximizing the area under the receiver operating characteristic curve (AUC) when only biased positive and unlabeled (PU) data are available. It leverages confidence scores—probabilities that an instance is positive—associated with a small set of labeled positives to derive an AUC risk estimator that accounts for bias. Experiments on eight real-world datasets demonstrate the method’s effectiveness.
Probabilistic models are typically trained using task-agnostic objectives like log-loss, which can lead to significant errors in downstream estimation. This disconnect is especially critical in Inverse Probability Weighting (IPW) for causal inference, where propensity score errors near $0$ and $1$ often lead to high bias and variance.
arXiv:2511. 22823v2 Announce Type: replace-cross Abstract: Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire.
The paper introduces a method for learning risk scores that remain reliable even when historical data contain unobserved confounders. By treating propensity weights as uncertain and applying sensitivity analysis with Wasserstein distributionally robust optimization, the authors formulate a robust learning problem solvable via an exponential cone program. Experiments on semi‑synthetic UCI data show the approach improves calibration by up to 29.2% over traditional benchmarks and 11.1% over the state of the art, without harming other performance metrics.
arXiv:2610.00284v1 Announce Type: cross Abstract: The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarize...
arXiv:2607. 11947v1 Announce Type: cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated.
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
arXiv:2608.30699v1 Announce Type: cross Abstract: Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to...
arXiv:2602. 17187v2 Announce Type: replace-cross Abstract: The problem of domain generalization concerns learning predictive models that are robust to distribution shifts when deployed in new, previously unseen environments.