arXiv Machine Learning By Leo Muxing Wang, Connor Mclaughlin, Lili Su

On the Power of Source Screening for Learning Shared Feature Extractors

Read the original on arXiv Machine Learning →

The paper investigates which data sources should be jointly learned when training shared feature extractors. Focusing on a linear setting where sources share a low‑dimensional subspace, it shows that carefully selecting a subset of sources—an informative subpopulation—can achieve minimax optimal subspace estimation, even when much data is discarded. The authors formalize this notion, propose algorithms and heuristics for identifying such subsets, and validate their effectiveness through theory and experiments on synthetic and real datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Transfer Learning for Matrix Completion

arXiv:2507.02248v2 Announce Type: replace-cross Abstract: In this paper, we explore the knowledge transfer under the setting of matrix completion, which aims to enhance the estimation of a low-rank t...

By Dali Liu, Yuying Xie, Haolei Weng
arXiv Machine Learning
Aug 31

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.

By Yiming Xie, Lili Su, Ningfang Mi
arXiv Machine Learning
6d ago

Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration

The paper compares two popular data‑integration techniques—Stack‑SVD, which concatenates datasets before performing singular value decomposition, and SVD‑Stack, which first decomposes each dataset separately and then aggregates the leading singular vectors. By deriving exact asymptotic performance expressions and phase transitions in a proportional regime, the authors show that neither method uniformly dominates the other when unweighted, but optimally weighted Stack‑SVD outperforms optimally weighted SVD‑Stack when the low‑rank signal is fully shared. They also demonstrate that SVD‑Stack can excel with partially shared components and provide practical algorithms for estimating optimal weights, supported by simulations and genomic experiments.

By Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry, Rong Ma