arXiv:2606. 06179v1 Announce Type: cross Abstract: Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
The paper introduces penalized nonreversible Langevin algorithms for sampling from a target distribution constrained to a compact convex set. It combines a squared distance penalty with skew-symmetric perturbations that preserve the penalized Gibbs distribution, and provides nonasymptotic total variation and Wasserstein bounds under various smoothness and contraction assumptions. Numerical experiments demonstrate the methods on constrained Bayesian regression, classification, neural networks, and truncated sampling, highlighting acceleration in a stochastic quadratic model.
By Pervez Ali, Weihao Dong, Xiaoyu Wang
arXiv:2608. 14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations.
By Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong
arXiv:2607. 20309v1 Announce Type: cross Abstract: Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions.
By William Kengne, Ehud Mossa Ockegna
arXiv:2409. 08469v4 Announce Type: replace-cross Abstract: We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\mathsf{KSD}$) and Wasserstein-2 metrics.
By Sayan Banerjee, Krishnakumar Balasubramanian, Promit Ghosal
arXiv:2609. 26647v1 Announce Type: cross Abstract: We study statistical rates in entropic optimal transport in the semi-discrete regime where one measure has finite support and the other is subGaussian.
By Tomas Gonzalez, Gonzalo Mena
arXiv:2606. 07325v1 Announce Type: cross Abstract: We study the minimax rate of estimating a future value $\mu_{t_n+h}$ of a curve $t\mapsto\mu_t$ in the $2$-Wasserstein space $\mathcal{P}_2(\mathbb{R}^d)$ from finitely many noisy snapshots of its past, under an adiabatic bound $\|\nabla_t^k v\|\le\varepsilon$ on the $k$-th covariant derivative of the velocity field.
By Munsik Kim
arXiv:2607. 29675v1 Announce Type: cross Abstract: Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored.
By Arkajyoti Bhattacharjee, Arnab Auddy
arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.
By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
arXiv:2509. 19830v3 Announce Type: replace Abstract: Kolmogorov-Arnold Networks (KANs) approximate multivariate functions by composing univariate transformations through additive or multiplicative aggregation.
By Wei Liu, Eleni Chatzi, Zhilu Lai
arXiv:2609. 25710v1 Announce Type: cross Abstract: The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data.
By Baicheng Li, Zuowei Shen, Haizhao Yang, Shijun Zhang
arXiv:2609. 12785v1 Announce Type: new Abstract: Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice.
By Misbah Uz Zaman, Anirbit Mukherjee