Hugging Face Trending Papers

Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness

We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.

arXiv AI
2d ago

Distributionally Robust Schr\"odinger Bridge

The paper introduces the Distributionally Robust Schr"odinger Bridge (DRSB), a method that learns a single controller capable of handling uncertainty in the initial distribution for stochastic transport tasks. DRSB’s objective combines control energy with a KL penalty on the terminal distribution, and it seeks to minimize the worst‑case value of this objective over an ambiguity set around the nominal initial distribution. The authors derive a variational formulation, connect it to stochastic optimal control and distributionally robust optimization, and propose an alternating algorithm with Wasserstein and Sinkhorn variants. Experiments on two‑dimensional transport and image‑to‑image translation demonstrate improved robustness to input perturbations compared to standard SB, while also achieving lower mean sliced Wasserstein distance on Gaussian mixture transport.

By Jinhwan Sul, Panagiotis Theodoropoulos, Vincent Pacelli, Jaemoo Choi, Evangelos Theodorou
arXiv Machine Learning
Jun 18

Generative models for decision-making under distributional shift

arXiv:2604. 04342v2 Announce Type: replace Abstract: Many data-driven decision problems are formulated using a nominal distribution estimated from historical data, while performance is ultimately determined by a deployment distribution that may be shifted, context-dependent, partially observed, or stress-induced.

By Xiuyuan Cheng, Yunqin Zhu, Yao Xie
arXiv Machine Learning
4d ago

Learning Distributionally Robust First-Order Methods for Convex Optimization

The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.

By Vinit Ranjan, Jisun Park, Bartolomeo Stellato