arXiv AI

Partial AUC Maximization from Positive-unlabeled Data

arXiv Machine Learning
Sep 11

AUC Maximization from Biased Positive-unlabeled Data with Confidence

The paper introduces a method for maximizing the area under the receiver operating characteristic curve (AUC) when only biased positive and unlabeled (PU) data are available. It leverages confidence scores—probabilities that an instance is positive—associated with a small set of labeled positives to derive an AUC risk estimator that accounts for bias. Experiments on eight real-world datasets demonstrate the method’s effectiveness.

By Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Kazuki Adachi, Yasuhiro Fujiwara
arXiv Machine Learning
Sep 22

Focused PU learning from imbalanced data

arXiv:2605.14467v2 Announce Type: replace Abstract: We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as dise...

By Elias Zavitsanos, Georgios Paliouras
arXiv Machine Learning
Jun 16

Imbalanced Classification under Capacity Constraints

arXiv:2605. 03289v2 Announce Type: replace-cross Abstract: Detecting observations from a minority class under severe class imbalance is a central challenge in applications such as fraud detection, medical screening, and industrial quality control.

By Daniel Fraiman, Ricardo Fraiman
arXiv Machine Learning
Sep 3

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.

By Ryota Ushio, Takashi Ishida, Masashi Sugiyama
arXiv AI
Sep 15

GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data

GRIN+ is a new machine unlearning framework that targets fast and precise data erasure in imbalanced medical datasets. It separates unlearning‑specific knowledge from general representations by analyzing gradient contributions of forget and retain sets, introduces a class‑adaptive influence scoring to counter gradient dominance, and uses a direction‑constrained update to protect essential clinical knowledge. Benchmarks on skin cancer, brain tumor, and breast ultrasound data show that GRIN+ balances privacy, efficiency, and utility, achieving high diagnostic accuracy and faster runtime than existing methods.

By Minghui Huang, Junxiao Wang