arXiv Machine Learning

Enhancing Conformal Prediction via Class Similarity

arXiv:2511. 19359v2 Announce Type: replace Abstract: Conformal Prediction (CP) has emerged as a powerful statistical framework for high-stakes classification applications.

arXiv Machine Learning
Sep 11

Benchmarking non-conformity score functions in conformal prediction

Conformal prediction replaces single-class predictions with prediction sets that guarantee a pre-specified coverage probability. The paper reviews properties of non‑conformity score functions, presents examples from the literature, and proposes new modifications. It introduces a method to evaluate prediction set sizes and compares different score functions, including their effectiveness for class‑conditional conformal prediction with imbalanced classes.

By Sol Erika Boman
arXiv Machine Learning
Sep 18

Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance

The paper introduces a method for combining heterogeneous, allied datasets—datasets that share the same class labels but have disjoint objects and largely distinct feature spaces—into a single unified feature space. By applying matrix completion to this merged space, the authors create a unified dataset that enables knowledge transfer between the original datasets. Experiments across multiple dataset pairs and classifiers show that models trained on the unified representation consistently outperform those trained separately on each dataset.

By Girish Keshav Palshikar
arXiv AI
Sep 1

CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding

CoLa-ICD is a knowledge‑enhanced framework designed to improve automatic medical coding of ICD codes in long, imbalanced clinical documents. It enriches ICD labels with external terms, models dependencies among related codes, and strengthens the alignment between label semantics and clinical evidence, particularly for rare codes. Experiments demonstrate that CoLa-ICD achieves state‑of‑the‑art performance in AUC, F1, and P@k, with larger gains in larger and sparser label spaces.

By Yihang Cheng, Veronica Liesaputra, Andrew Trotman
arXiv Machine Learning
Aug 28

MODIS: Multi-Omics Data Integration for Small and unpaired datasets

MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.

By Daniel Lepe-Soltero, Thierry Arti\`eres, Ana\"is Baudot, Paul Villoutreix
arXiv Machine Learning
Aug 4

xMICD: Explainable Representation of Multiple ICD Codes

arXiv:2608. 00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning.

By Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset
arXiv Machine Learning
Jul 31

What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis

arXiv:2510. 03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the machine learning community.

By Shahriar Kabir Nahin, Wenxiao Xiao, Joshua Liu, Anshuman Chhabra, Hongfu Liu
arXiv Machine Learning
Jul 9

A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It

arXiv:2607. 06605v1 Announce Type: new Abstract: Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns prediction sets containing the true label with probability at least 1 - alpha.

By Muhammadjon Tursunbadalov (School of Science and Technology, Champions College Prep, United States), Mustafojon Tursunbadalov (School of Science and Technology, Champions College Prep, United States)
arXiv AI
Jun 8

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

arXiv:2606. 07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns.

By Anurag Sharma, Sai Teja Chunchu, Prasenjit Mitra, Sandipan Sikdar, Koustav Rudra