arXiv:2607. 27975v1 Announce Type: new Abstract: We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions.
By Ka\u{g}an Akman, Naci Saldi, Serdar Y\"uksel
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
The paper introduces the Distributionally Robust Schr"odinger Bridge (DRSB), a method that learns a single controller capable of handling uncertainty in the initial distribution for stochastic transport tasks. DRSB’s objective combines control energy with a KL penalty on the terminal distribution, and it seeks to minimize the worst‑case value of this objective over an ambiguity set around the nominal initial distribution. The authors derive a variational formulation, connect it to stochastic optimal control and distributionally robust optimization, and propose an alternating algorithm with Wasserstein and Sinkhorn variants. Experiments on two‑dimensional transport and image‑to‑image translation demonstrate improved robustness to input perturbations compared to standard SB, while also achieving lower mean sliced Wasserstein distance on Gaussian mixture transport.
By Jinhwan Sul, Panagiotis Theodoropoulos, Vincent Pacelli, Jaemoo Choi, Evangelos Theodorou
arXiv:2609.23775v1 Announce Type: new
Abstract: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we...
By Suman Banerjee, Hiroyasu Tsukamoto
arXiv:2604. 04342v2 Announce Type: replace Abstract: Many data-driven decision problems are formulated using a nominal distribution estimated from historical data, while performance is ultimately determined by a deployment distribution that may be shifted, context-dependent, partially observed, or stress-induced.
By Xiuyuan Cheng, Yunqin Zhu, Yao Xie
arXiv:2606. 07600v1 Announce Type: cross Abstract: We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability measures.
By Albert Alcalde, Zhengping Ji, Enrique Zuazua
The paper introduces a distributionally robust method for learning hyperparameters of first‑order convex optimization algorithms. By minimizing a Wasserstein‑robust performance estimation problem over a dataset of problem instances, the approach interpolates between classical learning‑to‑optimize (L2O) and worst‑case PEP design. The authors solve the resulting problem with stochastic gradient descent, provide high‑probability risk bounds, and demonstrate that the learned algorithms outperform both worst‑case optimal and vanilla L2O baselines on logistic regression, LASSO, and linear programming tasks.
By Vinit Ranjan, Jisun Park, Bartolomeo Stellato
arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
By Yang Xu, Swetha Ganesh, Vaneet Aggarwal
arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2609.39512v1 Announce Type: new
Abstract: The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation a...
By Hong Zheng