arXiv Machine Learning

Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions

arXiv:2607. 26955v1 Announce Type: cross Abstract: Minimax-optimal rates for multivariate distribution estimation are known to suffer from the curse of dimensionality.

arXiv Machine Learning
Sep 21

Sparse Priors for Efficient Distribution Learning

arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.

By Saumya Goyal, Barnab\'as P\'oczos
arXiv AI
3d ago

BayesNDE: Bayesian Generative Modeling for Neural Density Estimation

BayesNDE is a neural density estimator that uses Bayesian generative modeling to estimate densities without relying on invertible networks or Jacobian-determinant calculations. It constructs an adaptive proposal for each observation by inferring a sample-specific latent posterior, and then applies bridge sampling to combine proposal samples with separate posterior samples for density estimation. Experiments on synthetic datasets show improved density estimation and structure recovery, while real-world applications demonstrate better anomaly detection.

By Chenglin Li, Qiao Liu
arXiv Machine Learning
Sep 25

Sufficiently Reduced Distributional Regression

Sufficiently Reduced Distributional Regression (SRDR) is a generative approach that merges conditional distribution estimation with nonlinear sufficient dimension reduction (SDR). By framing SDR as a risk minimization problem using strictly proper scoring rules, SRDR jointly learns a dimension reduction map and a generative prediction model through minimization of the energy score, which can be estimated via sampling. The method extends to multi‑environment data and classification, and theoretical results show convergence of estimated conditional distributions in energy distance, implying asymptotic sufficiency. In experiments on CT slice localization, superconductivity data, and digit classification, SRDR recovers low‑dimensional sufficient structure and matches or surpasses state‑of‑the‑art nonlinear SDR methods in representation quality and predictive performance.

By Alexander Henzi, Tiange Liu, Xinwei Shen
arXiv Machine Learning
Aug 28

A Flexible Empirical Bayes Approach to Generalized Linear Models, with Applications to Sparse Logistic Regression

The paper presents a tuning‑free empirical Bayes framework for Bayesian generalized linear models that uses a novel mean‑field variational inference algorithm. By estimating the prior within the VI procedure and optimizing the posterior mean directly, the method reduces optimization complexity and supports scalable solvers like L‑BFGS and stochastic gradient descent. Applied to sparse logistic regression, the approach shows superior predictive performance compared to existing methods.

By Dongyue Xie, Matthew Stephens
arXiv AI
Jun 9

Investigating the Histogram Loss in Regression

arXiv:2402. 13425v3 Announce Type: replace-cross Abstract: It is becoming increasingly common in regression to train neural networks that model the entire distribution even if only the mean is required for prediction.

By Ehsan Imani, Kai Luedemann, Sam Scholnick-Hughes, Esraa Elelimy, Martha White