arXiv AI

Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data

The paper proposes using the $q$Gaussian distribution, derived from Tsallis entropy maximization, to address the shortcomings of Gaussian assumptions in sparse learning with correlated and heterogeneous data. It introduces a new framework that adapts numerical equilibrium methods to composite optimization problems, applying it to the Hager‑Zhang conjugate gradient algorithm to create a stable, efficient sparse learning algorithm. The work offers both theoretical insights into alternative statistical distributions and practical tools for data analysis in fields like biostatistics.

arXiv Machine Learning
Aug 28

A Flexible Empirical Bayes Approach to Generalized Linear Models, with Applications to Sparse Logistic Regression

The paper presents a tuning‑free empirical Bayes framework for Bayesian generalized linear models that uses a novel mean‑field variational inference algorithm. By estimating the prior within the VI procedure and optimizing the posterior mean directly, the method reduces optimization complexity and supports scalable solvers like L‑BFGS and stochastic gradient descent. Applied to sparse logistic regression, the approach shows superior predictive performance compared to existing methods.

By Dongyue Xie, Matthew Stephens
arXiv Statistics ML
Sep 25

Riemannian Gradient Descent for Gaussian Mixture Models with unknown diagonal covariances

The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.

By Romane Giard, Yohann De Castro, Roland Denis, Cl\'ement Marteau
arXiv Machine Learning
Jun 9

Improving Bayesian Optimization via Training-Aware Conditional Diffusion Models

arXiv:2606. 08438v1 Announce Type: cross Abstract: Bayesian optimization (BO) is a widely used approach for black-box optimization that uses a Gaussian process (GP) as a surrogate and guides sequential evaluations via an acquisition function, with the ultimate goal of locating the global optimum $\mathbf{x}^{\star}$.

By Yilin Zheng, Haowei Wang, Szu Hui Ng, Enlu Zhou
arXiv Machine Learning
Aug 24

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

The paper introduces an amortized learning framework for selecting bandwidths in kernel density estimation by optimizing the logarithmic score across a distribution of tasks. It uses a truncated-and-renormalized bounded-support formulation and affine standardization to achieve stable learning and transferability across different intervals. Experiments on Gaussian samples, a multi-family benchmark, and randomized Gaussian mixtures demonstrate that the learned selector outperforms traditional methods such as Silverman’s rule, Sheather–Jones, and least‑squares cross‑validation, especially for small or heterogeneous samples.

By Junyi Liang, Hailiang Du
Hugging Face Trending Papers
Jul 15

Linear Independent Component Analysis via Optimal Transport

Linear Independent Component Analysis (ICA) recovers jointly independent source signals from their linear mixtures. To achieve this, classical ICA algorithms attempt to maximize non-Gaussianity, measured by negentropy, which is linked to independence by information theory.

arXiv Machine Learning
Sep 14

A Generalized Tangent Approximation based Variational Inference Framework for Strongly Super-Gaussian Likelihoods

The paper introduces a new variational inference framework that uses tangent transformations to handle strongly super‑Gaussian likelihoods across a wide range of probability models. By constructing tangent minorants of the log‑likelihood through convex duality, the method achieves conjugacy with Gaussian priors, enabling tractable inference where traditional approaches struggle. The authors provide algorithmic convergence guarantees and near‑parametric risk bounds, and demonstrate superior scalability and accuracy on both simulated and real‑world datasets compared to existing variational algorithms.

By Somjit Roy, Pritam Dey, Debdeep Pati, Bani K. Mallick