arXiv:2607. 03145v1 Announce Type: cross Abstract: The informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geometry of the data distribution.
By Alexandre L. M. Levada
The paper introduces a data‑driven method for learning Random Geometric Graphs (RGGs) in probabilistic metric spaces. It defines a distance function based on the cumulative distribution of a disparity variable that captures differences in vertex connectivity and correlation of attached random variables, enabling edges to exist with a specified probability. The approach includes a rejection‑sampling technique for edge probability estimation and a closed‑form posterior for learning the inter‑observable correlation matrix, and it is demonstrated on highly multivariate real datasets.
By Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu
arXiv:2607. 06644v1 Announce Type: cross Abstract: Determinantal point processes have recently emerged as a kernel-based alternative to standard independent sampling for constructing efficient minibatches, coresets, and other compact representations of large-scale datasets.
By Hoang-Son Tran, Pranav Gupta, Subhroshekhar Ghosh
arXiv:2409. 18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond.
By Iskander Azangulov, George Deligiannidis, Judith Rousseau
arXiv:2609.31458v1 Announce Type: new
Abstract: Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large languag...
By Jaehee Seo, Jisu Kim
arXiv:2606. 14334v1 Announce Type: new Abstract: High-dimensional datasets often concentrate near low-dimensional structures, but estimating their geometry from samples typically relies on graphs and kernels that scale poorly with dataset size and dimension.
By Jacob Bamberger, Adam Gosztolai, Pierre Vandergheynst, Michael Bronstein, Iolo Jones