arXiv Machine Learning

Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance

The paper introduces a method for combining heterogeneous, allied datasets—datasets that share the same class labels but have disjoint objects and largely distinct feature spaces—into a single unified feature space. By applying matrix completion to this merged space, the authors create a unified dataset that enables knowledge transfer between the original datasets. Experiments across multiple dataset pairs and classifiers show that models trained on the unified representation consistently outperform those trained separately on each dataset.

arXiv Machine Learning
Sep 22

Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation

The paper introduces FedPFT, a federated learning framework that tackles the feature‑classifier mismatch problem by using personalized prompts processed through a shared self‑attention transformation module. Unlike prior methods that either degrade the feature extractor or address the mismatch only after training, FedPFT aligns local features with the global classifier during training, improving aggregation and model performance. Experiments demonstrate that FedPFT surpasses state‑of‑the‑art methods by up to 5.07%, and gains up to 7.08% when combined with collaborative contrastive learning.

By Xinghao Wu, Xuefeng Liu, Jianwei Niu, Guogang Zhu, Mingjia Shi, Shaojie Tang, Jing Yuan
arXiv Machine Learning
Aug 11

Coupled Training with Privileged Information and Unlabeled Data

arXiv:2605. 23268v2 Announce Type: replace-cross Abstract: In many prediction problems, we have extra information during training (for example, measurements that are expensive or slow to collect) that will not be available when the model is deployed.

By Jiahao Shi, Omar Hagrass, Jason M. Klusowski
arXiv Machine Learning
Aug 31

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.

By Yiming Xie, Lili Su, Ningfang Mi
arXiv Machine Learning
Sep 14

Class-wise Contribution Estimation via Logit Maximization for Federated Learning

The paper introduces CELM, a data‑free framework for federated learning that estimates class‑wise contribution by maximizing logits. It constructs a cross‑client evidence matrix to quantify each client’s competence and coverage for each class, then uses this matrix to compute weighted aggregation that upweights clients offering strong evidence for underrepresented classes. The method maintains stability through simplex constraints and momentum smoothing, and it is compatible with standard FL pipelines, showing improved robustness to class imbalance and heterogeneity on vision benchmarks.

By Asim Ukaye, Nurbek Tastan, Mubarak Abdu-Aguye, Karthik Nandakumar
arXiv Machine Learning
Jun 17

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.

By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv Machine Learning
Aug 18

FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting

arXiv:2608. 15310v1 Announce Type: cross Abstract: Multimodal data collected by heterogeneous devices are used for collaborative training, where federated learning (FL) serves as a key paradigm for effective distributed modeling with data privacy preservation.

By Zhenyan Liu, Hua Zhang, Haoran Gao, Qi Li, Hongliang Zhu, Huiyu Zhou, Zongliang Shen, Yanxin Xu, Jiahui Wang
arXiv Machine Learning
Aug 28

MODIS: Multi-Omics Data Integration for Small and unpaired datasets

MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.

By Daniel Lepe-Soltero, Thierry Arti\`eres, Ana\"is Baudot, Paul Villoutreix
arXiv Machine Learning
Sep 24

Binary Classification from Coupled Pairwise Labels

The paper introduces SD-Pcomp learning, a binary classification framework that jointly utilizes Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels from instance pairs. It proposes an objective function that can be decomposed into either an SD estimator plus ordering information or a Pcomp estimator plus pair-type information, thereby integrating complementary relational cues. Experiments on eight datasets demonstrate that combining both label types improves classification accuracy and AUC compared to using either alone or a simple convex combination.

By Tomoya Tate, Kosuke Sugiyama, Masato Uchida