arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
The paper investigates how policy learning algorithms should balance expected welfare against sampling risk in evidence-based policymaking. It demonstrates that algorithmic stability—specifically, a policy’s insensitivity to the replacement of a single experimental unit—limits sampling risk. The authors introduce policy‑vote bagging, which trains on many subsamples and averages their votes, preserving expected welfare while improving expected utility for risk‑averse researchers, and provide sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.
By Harvey Barnhard, Giacomo Opocher, Rahul Singh
arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.
By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.
By Yiming Xie, Lili Su, Ningfang Mi
arXiv:2606. 30372v1 Announce Type: new Abstract: Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias.
By Haobo Yang
The paper investigates regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes, a common issue in fields such as clinical trials, labor markets, and auctions. It provides a complete characterization of the minimal assumptions on the functional forms of selection processes that allow regression to remain possible, and shows that the regression function can sometimes be identified even when the selection filter itself cannot. Under stronger identification conditions, the authors also deliver finite‑sample estimation guarantees, explicit convergence rates, and oracle‑efficient algorithms, offering the first general‑purpose estimation method for this broad class of selection problems.
By Vikram Kher, Jane H. Lee, Anay Mehrotra, Manolis Zampetakis
arXiv:2506. 20573v4 Announce Type: replace-cross Abstract: Public datasets, crucial for modern machine learning and statistical inference, often contain low-quality or contaminated samples that can harm model performance.
By Kristian Minchev, Dimitar I. Dimitrov, Nikola Konstantinov
The paper investigates which data sources should be jointly learned when training shared feature extractors. Focusing on a linear setting where sources share a low‑dimensional subspace, it shows that carefully selecting a subset of sources—an informative subpopulation—can achieve minimax optimal subspace estimation, even when much data is discarded. The authors formalize this notion, propose algorithms and heuristics for identifying such subsets, and validate their effectiveness through theory and experiments on synthetic and real datasets.
By Leo Muxing Wang, Connor Mclaughlin, Lili Su
The paper introduces a bias‑correction method for synthetic oversampling in imbalanced learning. It estimates the loss discrepancy caused by the data generator using a held‑out majority subset and transfers this correction to the minority class under a uniform bias‑transfer assumption. The authors provide finite‑sample bounds for bias transfer and excess balanced risk, identify when SMOTE introduces significant bias, and demonstrate the method’s applicability to multi‑task learning and propensity‑score estimation, with empirical results showing greatest benefit when synthetic distortion is large.
By Pengfei Lyu, Zhengchi Ma, Linjun Zhang, Anru R. Zhang
The paper introduces the first federated algorithm for Targeted Maximum Likelihood Estimation (TMLE), enabling hospitals, banks, or registries to perform TMLE without sharing individual data. Two frameworks—FedTMLE‑G, which aggregates local gradients, and FedTMLE‑L, which allows each institution to complete its own fluctuation fit—are presented, along with a finite‑precision communication protocol that keeps numerical targeting error negligible. The authors also discuss privacy implications, convergence bounds, and trade‑offs between institutional influence and sampling variability.
By Diyang Li, Fei Wang, Kyra Gan
arXiv:2608. 10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute $S$ requires a representation $Z$ that is statistically independent of $S$.
By Yijin Ni, Xiaoming Huo
arXiv:2609.06294v1 Announce Type: new
Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental s...
By Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez, Tamar Krishnamurti, Bryan Wilder