arXiv Machine Learning By Gilles Blanchard (LMO, DATASHAPE), Jean-Baptiste Fermanian (LMO), Hannah Marienwald (TUB)

Estimation of multiple mean vectors in high dimension

Read the original on arXiv Machine Learning →

The paper proposes methods for estimating many high‑dimensional mean vectors from independent samples by forming convex combinations of empirical means. Two data‑dependent weighting strategies are introduced: one uses a testing procedure to pick low‑variance neighbouring means, yielding a closed‑form plug‑in formula; the other minimizes an upper confidence bound on quadratic risk. Theoretical results show these approaches asymptotically achieve oracle (minimax) risk improvements as the effective dimension grows, and experiments confirm their effectiveness on simulated and real kernel mean embedding tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 28

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.

By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv Machine Learning
Sep 25

Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration

The paper compares two popular data‑integration techniques—Stack‑SVD, which concatenates datasets before performing singular value decomposition, and SVD‑Stack, which first decomposes each dataset separately and then aggregates the leading singular vectors. By deriving exact asymptotic performance expressions and phase transitions in a proportional regime, the authors show that neither method uniformly dominates the other when unweighted, but optimally weighted Stack‑SVD outperforms optimally weighted SVD‑Stack when the low‑rank signal is fully shared. They also demonstrate that SVD‑Stack can excel with partially shared components and provide practical algorithms for estimating optimal weights, supported by simulations and genomic experiments.

By Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry, Rong Ma
arXiv Machine Learning
Jul 7

A Gradient Flow Perspective on Minimum MMD Estimation

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
Hugging Face Trending Papers
Sep 24

On the SoS Certifiability of Log-Concave Distributions

For an arbitrary isotropic log-concave distribution $P$ on $\mathbb{R}^d$, we prove that the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares for every even $m\ge2$, where $C>0$ is a universal constant. This removes the dependence on the Poincaré constant in the theorem of Kothari and Steinhardt (arXiv:1711.