arXiv:2604. 07635v2 Announce Type: replace-cross Abstract: This research considers a scalable inference for spatial data modeled through Gaussian intrinsic conditional autoregressive (ICAR) structures.
By Debjoy Thakur
The paper introduces localized diffusion models, which exploit locality structure—sparse conditional dependencies among target variables—to reduce the dimensionality of the score function. By training a localized neural network with a localized score matching loss, the authors demonstrate that diffusion models can achieve dimension‑independent error bounds, balancing statistical and localization errors with a moderate radius. This approach also enables parallel training, potentially improving efficiency for large‑scale applications.
By Georg A. Gottwald, Shuigen Liu, Youssef Marzouk, Sebastian Reich, Xin T. Tong
arXiv:2607. 06644v1 Announce Type: cross Abstract: Determinantal point processes have recently emerged as a kernel-based alternative to standard independent sampling for constructing efficient minibatches, coresets, and other compact representations of large-scale datasets.
By Hoang-Son Tran, Pranav Gupta, Subhroshekhar Ghosh
arXiv:2607. 20502v1 Announce Type: new Abstract: To allow for principled comparison between two probabilistic graphical models defined over non-identical variable sets, they have to be lifted to a common measurable space.
By Jan Speller, Malte Luttermann, Marcel Gehrke, Tanya Braun
arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.
By Saumya Goyal, Barnab\'as P\'oczos
arXiv:2406. 12659v3 Announce Type: replace-cross Abstract: We propose a scalable variational Bayes method for statistical inference for a single or pre-specified low-dimensional subset of the coordinates of a high-dimensional parameter in sparse linear regression.
By Isma\"el Castillo, Alice L'Huillier, Kolyan Ray, Luke Travis
The paper introduces linear‑time randomized algorithms for unbiased approximation of general random walk kernels (RWKs) on sparse graphs, covering both labelled and unlabelled cases. By sampling dependent random walks and constructing novel graph embeddings in ρ^d, the method avoids building the direct product graph, enabling scaling to massive datasets that cannot fit on a single machine. The authors provide exponential concentration bounds for the estimator’s sharpness and demonstrate up to 27× speed‑ups and 128× larger graph handling compared to previous cubic‑time approaches.
By Krzysztof Choromanski, Isaac Reid, Arijit Sehanobish, Avinava Dubey
arXiv:2607. 28706v1 Announce Type: cross Abstract: We study the convergence properties of the random-sweep Gibbs sampler for Gaussian graphical models with a thin-membrane prior.
By Borna Khodabandeh, Mehdi Molkaraie
arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2511. 08307v2 Announce Type: replace-cross Abstract: Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries.
By Aranyak Acharyya, Joshua Agterberg, Youngser Park, Carey E. Priebe
The paper reviews methods for adding a new point to a vector diagram using proximity data, a problem first examined by J.C. Gower in 1968. It classifies existing kernel-based approaches into two strategies: projection, analogous to adding a point in principal component analysis, and restricted reconstruction, which seeks to re‑optimize the multivariate analysis while keeping the existing diagram fixed. The authors show that each method can be derived from one of these two strategies and discuss when each strategy may be preferable.
By Michael W. Trosset, Kaiyi Tan, Minh Tang, Carey E. Priebe
arXiv:2608. 15121v1 Announce Type: cross Abstract: Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the central subspace (CS).
By Ye Tian