We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem.
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
The paper introduces the Distributionally Robust Schr"odinger Bridge (DRSB), a method that learns a single controller capable of handling uncertainty in the initial distribution for stochastic transport tasks. DRSB’s objective combines control energy with a KL penalty on the terminal distribution, and it seeks to minimize the worst‑case value of this objective over an ambiguity set around the nominal initial distribution. The authors derive a variational formulation, connect it to stochastic optimal control and distributionally robust optimization, and propose an alternating algorithm with Wasserstein and Sinkhorn variants. Experiments on two‑dimensional transport and image‑to‑image translation demonstrate improved robustness to input perturbations compared to standard SB, while also achieving lower mean sliced Wasserstein distance on Gaussian mixture transport.
By Jinhwan Sul, Panagiotis Theodoropoulos, Vincent Pacelli, Jaemoo Choi, Evangelos Theodorou
arXiv:2609.23775v1 Announce Type: new
Abstract: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we...
By Suman Banerjee, Hiroyasu Tsukamoto
arXiv:2604. 04342v2 Announce Type: replace Abstract: Many data-driven decision problems are formulated using a nominal distribution estimated from historical data, while performance is ultimately determined by a deployment distribution that may be shifted, context-dependent, partially observed, or stress-induced.
By Xiuyuan Cheng, Yunqin Zhu, Yao Xie
arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang