arXiv:2607. 27975v1 Announce Type: new Abstract: We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions.
By Ka\u{g}an Akman, Naci Saldi, Serdar Y\"uksel
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2604. 04342v2 Announce Type: replace Abstract: Many data-driven decision problems are formulated using a nominal distribution estimated from historical data, while performance is ultimately determined by a deployment distribution that may be shifted, context-dependent, partially observed, or stress-induced.
By Xiuyuan Cheng, Yunqin Zhu, Yao Xie
arXiv:2606. 07600v1 Announce Type: cross Abstract: We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability measures.
By Albert Alcalde, Zhengping Ji, Enrique Zuazua
arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
By Yang Xu, Swetha Ganesh, Vaneet Aggarwal