arXiv Machine Learning By Simon Chung, Colby J. Vorland, Donna L. Maney, Andrew W. Brown

A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research

Read the original on arXiv Machine Learning →

The paper introduces a sampling algorithm based on a multivariate Bernoulli distribution to address challenges in multi‑label datasets where labels are non‑exclusive and vary widely in frequency. By estimating distribution parameters from observed label frequencies and computing weights for each label combination, the method produces weighted samples that reflect a target distribution while respecting label dependencies. Applied to Web of Science research articles labeled with 64 biomedical topics, the approach yielded a more balanced sub‑sample, improving representation of minority categories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 28

Robust Conformalized Selection with Noisy Responses

arXiv:2607. 22985v1 Announce Type: cross Abstract: Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models.

By Chengyao Yu, Hongxin Wei, Bingyi Jing
arXiv Machine Learning
1d ago

How many labelers do you have? A closer look at gold-standard labels

The paper examines the common practice of aggregating multiple labels per instance into a single ‘true’ label for supervised learning. By creating a theoretical model, the authors show that using the full, non‑aggregated label information can make it easier to train well‑calibrated models, though the benefits depend on the specific problem. They predict when non‑aggregated labels will improve learning and validate these predictions on real datasets.

By Chen Cheng, Hilal Asi, John Duchi