arXiv Machine Learning

Learning from Biased and Costly Data Sources: Minimax-optimal Data Collection under a Budget

arXiv:2602. 17894v2 Announce Type: replace-cross Abstract: Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest.

arXiv Machine Learning
Jun 26

Learning from a Biased Sample

arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.

By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv Machine Learning
Sep 18

Stable Policy Learning

The paper investigates how policy learning algorithms should balance expected welfare against sampling risk in evidence-based policymaking. It demonstrates that algorithmic stability—specifically, a policy’s insensitivity to the replacement of a single experimental unit—limits sampling risk. The authors introduce policy‑vote bagging, which trains on many subsamples and averages their votes, preserving expected welfare while improving expected utility for risk‑averse researchers, and provide sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.

By Harvey Barnhard, Giacomo Opocher, Rahul Singh
arXiv Machine Learning
Jun 15

A Complexity Measure for Active Learning in Multi-group Mean Estimation

arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.

By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
arXiv Machine Learning
Aug 31

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.

By Yiming Xie, Lili Su, Ningfang Mi
arXiv Statistics ML
3d ago

Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases

The paper investigates regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes, a common issue in fields such as clinical trials, labor markets, and auctions. It provides a complete characterization of the minimal assumptions on the functional forms of selection processes that allow regression to remain possible, and shows that the regression function can sometimes be identified even when the selection filter itself cannot. Under stronger identification conditions, the authors also deliver finite‑sample estimation guarantees, explicit convergence rates, and oracle‑efficient algorithms, offering the first general‑purpose estimation method for this broad class of selection problems.

By Vikram Kher, Jane H. Lee, Anay Mehrotra, Manolis Zampetakis
arXiv Machine Learning
Jun 9

LARP: Learner-Agnostic Robust Data Prefiltering

arXiv:2506. 20573v4 Announce Type: replace-cross Abstract: Public datasets, crucial for modern machine learning and statistical inference, often contain low-quality or contaminated samples that can harm model performance.

By Kristian Minchev, Dimitar I. Dimitrov, Nikola Konstantinov
arXiv Machine Learning
Aug 20

On the Power of Source Screening for Learning Shared Feature Extractors

The paper investigates which data sources should be jointly learned when training shared feature extractors. Focusing on a linear setting where sources share a low‑dimensional subspace, it shows that carefully selecting a subset of sources—an informative subpopulation—can achieve minimax optimal subspace estimation, even when much data is discarded. The authors formalize this notion, propose algorithms and heuristics for identifying such subsets, and validate their effectiveness through theory and experiments on synthetic and real datasets.

By Leo Muxing Wang, Connor Mclaughlin, Lili Su
arXiv Machine Learning
Sep 14

Bias-Corrected Data Synthesis for Imbalanced Learning

The paper introduces a bias‑correction method for synthetic oversampling in imbalanced learning. It estimates the loss discrepancy caused by the data generator using a held‑out majority subset and transfers this correction to the minority class under a uniform bias‑transfer assumption. The authors provide finite‑sample bounds for bias transfer and excess balanced risk, identify when SMOTE introduces significant bias, and demonstrate the method’s applicability to multi‑task learning and propensity‑score estimation, with empirical results showing greatest benefit when synthetic distortion is large.

By Pengfei Lyu, Zhengchi Ma, Linjun Zhang, Anru R. Zhang
arXiv Machine Learning
5d ago

Federated Targeted Maximum Likelihood Estimation

The paper introduces the first federated algorithm for Targeted Maximum Likelihood Estimation (TMLE), enabling hospitals, banks, or registries to perform TMLE without sharing individual data. Two frameworks—FedTMLE‑G, which aggregates local gradients, and FedTMLE‑L, which allows each institution to complete its own fluctuation fit—are presented, along with a finite‑precision communication protocol that keeps numerical targeting error negligible. The authors also discuss privacy implications, convergence bounds, and trade‑offs between institutional influence and sampling variability.

By Diyang Li, Fei Wang, Kyra Gan