arXiv Machine Learning

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

arXiv:2608. 13418v1 Announce Type: cross Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution.

Hugging Face Trending Papers
Aug 13

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data.

arXiv AI
1d ago

Distributionally Robust Survival Models under Subpopulation Shift and Outlier Contamination

The paper introduces a distributionally robust framework for survival analysis that simultaneously tackles latent subpopulation shift and outlier contamination. It employs an outer minimization to refine the nominal distribution by down-weighting contaminated samples and an inner maximization to target the most challenging subpopulation, directly handling non-decomposable survival losses such as the Cox partial log-likelihood. An alternating gradient-based algorithm, guided by KKT conditions, is developed, and experiments on simulated data and two benchmarks show improved worst-group performance and stable training under contamination.

By Seonghwi Kim, Sung Ho Jo, Minwoo Chae
arXiv Machine Learning
Aug 26

Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling

The paper introduces MS‑WDRO, a multi‑source Wasserstein distributionally robust optimization framework for reconstructing complex network topologies from scarce target‑domain data and abundant heterogeneous source data. It fuses sources via a weighted Wasserstein barycenter, builds an ambiguity set around it, and solves a regularized Laplacian estimator using a provably convergent ADMM scheme. The authors provide finite‑sample guarantees, demonstrate that naive aggregation is suboptimal, and show through experiments on synthetic data and the ABIDE I neuroimaging dataset that MS‑WDRO outperforms seven baselines in graph recovery, sample efficiency, and diagnostic utility, especially when target samples are limited.

By Chuansen Peng, Yifan Xia, Jinshan Zhong, Xiaojing Shen
arXiv Machine Learning
6d ago

Learning Distributionally Robust First-Order Methods for Convex Optimization

The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.

By Vinit Ranjan, Jisun Park, Bartolomeo Stellato