arXiv Statistics ML

Stein's method for marginals on large graphical models

arXiv Machine Learning
Sep 24

Localized Diffusion Models

The paper introduces localized diffusion models, which exploit locality structure—sparse conditional dependencies among target variables—to reduce the dimensionality of the score function. By training a localized neural network with a localized score matching loss, the authors demonstrate that diffusion models can achieve dimension‑independent error bounds, balancing statistical and localization errors with a moderate radius. This approach also enables parallel training, potentially improving efficiency for large‑scale applications.

By Georg A. Gottwald, Shuigen Liu, Youssef Marzouk, Sebastian Reich, Xin T. Tong
arXiv Machine Learning
Sep 21

Sparse Priors for Efficient Distribution Learning

arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.

By Saumya Goyal, Barnab\'as P\'oczos
arXiv Machine Learning
Aug 27

Optimal Time Complexity Algorithms for Computing General Random Walk Graph Kernels on Sparse Graphs

The paper introduces linear‑time randomized algorithms for unbiased approximation of general random walk kernels (RWKs) on sparse graphs, covering both labelled and unlabelled cases. By sampling dependent random walks and constructing novel graph embeddings in ρ^d, the method avoids building the direct product graph, enabling scaling to massive datasets that cannot fit on a single machine. The authors provide exponential concentration bounds for the estimator’s sharpness and demonstrate up to 27× speed‑ups and 128× larger graph handling compared to previous cubic‑time approaches.

By Krzysztof Choromanski, Isaac Reid, Arijit Sehanobish, Avinava Dubey
arXiv Machine Learning
Sep 18

Out-of-Sample Embedding with Proximity Data: Projection versus Restricted Reconstruction

The paper reviews methods for adding a new point to a vector diagram using proximity data, a problem first examined by J.C. Gower in 1968. It classifies existing kernel-based approaches into two strategies: projection, analogous to adding a point in principal component analysis, and restricted reconstruction, which seeks to re‑optimize the multivariate analysis while keeping the existing diagram fixed. The authors show that each method can be derived from one of these two strategies and discuss when each strategy may be preferable.

By Michael W. Trosset, Kaiyi Tan, Minh Tang, Carey E. Priebe