Conformal Prediction with Macro-Coverage Guarantees
arXiv:2606. 28598v1 Announce Type: cross Abstract: Prediction sets should have high coverage to be useful, but some coverage notions are more practically relevant than others.
arXiv:2511. 19359v2 Announce Type: replace Abstract: Conformal Prediction (CP) has emerged as a powerful statistical framework for high-stakes classification applications.
arXiv:2606. 28598v1 Announce Type: cross Abstract: Prediction sets should have high coverage to be useful, but some coverage notions are more practically relevant than others.
Conformal prediction replaces single-class predictions with prediction sets that guarantee a pre-specified coverage probability. The paper reviews properties of non‑conformity score functions, presents examples from the literature, and proposes new modifications. It introduces a method to evaluate prediction set sizes and compares different score functions, including their effectiveness for class‑conditional conformal prediction with imbalanced classes.
The paper introduces a method for combining heterogeneous, allied datasets—datasets that share the same class labels but have disjoint objects and largely distinct feature spaces—into a single unified feature space. By applying matrix completion to this merged space, the authors create a unified dataset that enables knowledge transfer between the original datasets. Experiments across multiple dataset pairs and classifiers show that models trained on the unified representation consistently outperform those trained separately on each dataset.
CoLa-ICD is a knowledge‑enhanced framework designed to improve automatic medical coding of ICD codes in long, imbalanced clinical documents. It enriches ICD labels with external terms, models dependencies among related codes, and strengthens the alignment between label semantics and clinical evidence, particularly for rare codes. Experiments demonstrate that CoLa-ICD achieves state‑of‑the‑art performance in AUC, F1, and P@k, with larger gains in larger and sparser label spaces.
Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long documents, imbalanced label distributions, and diverse terms. These challenges are especially sever...
arXiv:2606. 31577v1 Announce Type: cross Abstract: Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees.
arXiv:2606. 14909v1 Announce Type: cross Abstract: We consider the problem of uncertainty quantification for a pretrained classification model deployed under unknown distribution shift.
MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.
arXiv:2608. 00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning.
arXiv:2510. 03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the machine learning community.
arXiv:2607. 06605v1 Announce Type: new Abstract: Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns prediction sets containing the true label with probability at least 1 - alpha.
arXiv:2606. 07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns.