arXiv Machine Learning

A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research

The paper introduces a sampling algorithm based on a multivariate Bernoulli distribution to address challenges in multi‑label datasets where labels are non‑exclusive and vary widely in frequency. By estimating distribution parameters from observed label frequencies and computing weights for each label combination, the method produces weighted samples that reflect a target distribution while respecting label dependencies. Applied to Web of Science research articles labeled with 64 biomedical topics, the approach yielded a more balanced sub‑sample, improving representation of minority categories.

arXiv Machine Learning
Jul 28

Robust Conformalized Selection with Noisy Responses

arXiv:2607. 22985v1 Announce Type: cross Abstract: Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models.

By Chengyao Yu, Hongxin Wei, Bingyi Jing
arXiv Machine Learning
1d ago

How many labelers do you have? A closer look at gold-standard labels

The paper examines the common practice of aggregating multiple labels per instance into a single ‘true’ label for supervised learning. By creating a theoretical model, the authors show that using the full, non‑aggregated label information can make it easier to train well‑calibrated models, though the benefits depend on the specific problem. They predict when non‑aggregated labels will improve learning and validate these predictions on real datasets.

By Chen Cheng, Hilal Asi, John Duchi
arXiv AI
Jun 8

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

arXiv:2606. 07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns.

By Anurag Sharma, Sai Teja Chunchu, Prasenjit Mitra, Sandipan Sikdar, Koustav Rudra
arXiv Machine Learning
Jul 8

Factorizable joint shift revisited

arXiv:2601. 15036v4 Announce Type: replace Abstract: Factorizable joint shift (FJS) represents a type of distribution shift (or dataset shift) that comprises both covariate and label shift.

By Dirk Tasche
arXiv Machine Learning
Jul 21

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

arXiv:2607. 17653v1 Announce Type: cross Abstract: Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data.

By Jing Li, Pan Liu, Meng Zhao, Wanli Xue, Yanhong Yang, Xu Cheng, Fan Shi, Jianhua Zhang, Qinghua Hu, Shengyong Chen