Not all Jensen-Shannon Divergence Estimators are Equal
arXiv:2606. 16411v1 Announce Type: new Abstract: The Jensen-Shannon divergence is widely reported as a scalar measure of fidelity for synthetic tabular data.
arXiv:2601. 22784v2 Announce Type: replace-cross Abstract: We introduce a rank-statistic approximation of $f$-divergences that avoids explicit density-ratio estimation by working directly with the distribution of ranks.
arXiv:2606. 16411v1 Announce Type: new Abstract: The Jensen-Shannon divergence is widely reported as a scalar measure of fidelity for synthetic tabular data.
arXiv:2609. 38412v1 Announce Type: cross Abstract: We propose a novel local-polynomial estimator of the ratio $r=f/g$ of two $d$-dimensional densities $f$ and $g$, from which independent samples are available.
arXiv:2606. 30310v1 Announce Type: cross Abstract: The Sliced Wasserstein (SW) distance has emerged as a computationally attractive alternative to the Wasserstein distance by leveraging one-dimensional optimal transport along random projections.
arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.
arXiv:2606. 16257v1 Announce Type: cross Abstract: Sampling from high-dimensional, non-log-concave distributions with unnormalized densities is a fundamental challenge in machine learning, particularly when the exact gradient of the potential is unavailable and must be approximated via stochastic gradients that exhibit high variance under a fixed budget of gradient computations per iteration.
The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.
arXiv:2608. 13922v1 Announce Type: new Abstract: Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified.
arXiv:2607. 20119v1 Announce Type: cross Abstract: We introduce the Directional Kernel Mean Difference (DKMD), a signed statistic for univariate distribution comparison that preserves the direction of distributional shifts.
The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.
arXiv:2609.27546v1 Announce Type: cross Abstract: Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution...
arXiv:2605. 26000v2 Announce Type: replace-cross Abstract: Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization.
arXiv:2609. 20749v1 Announce Type: cross Abstract: Location estimation exhibits markedly different finite-sample behavior across noise distributions: regular families typically yield root-\(n\) rates, whereas compactly supported laws may admit faster, boundary-driven rates.