Enhancing Conformal Prediction via Class Similarity
arXiv:2511. 19359v2 Announce Type: replace Abstract: Conformal Prediction (CP) has emerged as a powerful statistical framework for high-stakes classification applications.
The paper introduces a method for combining heterogeneous, allied datasets—datasets that share the same class labels but have disjoint objects and largely distinct feature spaces—into a single unified feature space. By applying matrix completion to this merged space, the authors create a unified dataset that enables knowledge transfer between the original datasets. Experiments across multiple dataset pairs and classifiers show that models trained on the unified representation consistently outperform those trained separately on each dataset.
arXiv:2511. 19359v2 Announce Type: replace Abstract: Conformal Prediction (CP) has emerged as a powerful statistical framework for high-stakes classification applications.
The paper introduces FedPFT, a federated learning framework that tackles the feature‑classifier mismatch problem by using personalized prompts processed through a shared self‑attention transformation module. Unlike prior methods that either degrade the feature extractor or address the mismatch only after training, FedPFT aligns local features with the global classifier during training, improving aggregation and model performance. Experiments demonstrate that FedPFT surpasses state‑of‑the‑art methods by up to 5.07%, and gains up to 7.08% when combined with collaborative contrastive learning.
arXiv:2605. 23268v2 Announce Type: replace-cross Abstract: In many prediction problems, we have extra information during training (for example, measurements that are expensive or slow to collect) that will not be available when the model is deployed.
arXiv:2607. 24943v1 Announce Type: cross Abstract: In many classification problems, reliable instance-level labels are unavailable.
The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.
arXiv:2606. 10250v1 Announce Type: cross Abstract: Class imbalance is a common problem in deep learning that severely degrades performance.
arXiv:2606. 11616v1 Announce Type: new Abstract: High-quality training data is essential for the success of machine learning models.
The paper introduces CELM, a data‑free framework for federated learning that estimates class‑wise contribution by maximizing logits. It constructs a cross‑client evidence matrix to quantify each client’s competence and coverage for each class, then uses this matrix to compute weighted aggregation that upweights clients offering strong evidence for underrepresented classes. The method maintains stability through simplex constraints and momentum smoothing, and it is compatible with standard FL pipelines, showing improved robustness to class imbalance and heterogeneity on vision benchmarks.
arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
arXiv:2608. 15310v1 Announce Type: cross Abstract: Multimodal data collected by heterogeneous devices are used for collaborative training, where federated learning (FL) serves as a key paradigm for effective distributed modeling with data privacy preservation.
MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.
The paper introduces SD-Pcomp learning, a binary classification framework that jointly utilizes Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels from instance pairs. It proposes an objective function that can be decomposed into either an SD estimator plus ordering information or a Pcomp estimator plus pair-type information, thereby integrating complementary relational cues. Experiments on eight datasets demonstrate that combining both label types improves classification accuracy and AUC compared to using either alone or a simple convex combination.