arXiv Machine Learning By Kisung You, Boram Cho

HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces

Read the original on arXiv Machine Learning →

arXiv:2607. 27532v1 Announce Type: cross Abstract: Heavy tails weaken high-confidence control for the empirical mean.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 3

Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle

The paper revisits median‑of‑means estimation from a deterministic optimization perspective, introducing a family of block‑Lp estimators (for 0 < p ≤ 1) that achieve robust learning with heavy‑tailed and adversarially corrupted data. It shows that any convex block M‑estimator cannot attain the trimmed‑block oracle constant, while the nonconvex block‑Lp family provides finite‑sample robustness bounds that approach this oracle constant as p decreases. The authors also prove that the block‑Lp objectives have a benign landscape—every local minimum is close to the true parameter—and combine these results with block‑level concentration to obtain sub‑Gaussian deviation bounds under finite 2+δ moments, extending to high‑dimensional robust mean estimation and sparse regression.

By Angshul Majumdar
arXiv Machine Learning
Jul 28

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.

By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv Machine Learning
Sep 18

Estimation of multiple mean vectors in high dimension

The paper proposes methods for estimating many high‑dimensional mean vectors from independent samples by forming convex combinations of empirical means. Two data‑dependent weighting strategies are introduced: one uses a testing procedure to pick low‑variance neighbouring means, yielding a closed‑form plug‑in formula; the other minimizes an upper confidence bound on quadratic risk. Theoretical results show these approaches asymptotically achieve oracle (minimax) risk improvements as the effective dimension grows, and experiments confirm their effectiveness on simulated and real kernel mean embedding tasks.

By Gilles Blanchard (LMO, DATASHAPE), Jean-Baptiste Fermanian (LMO), Hannah Marienwald (TUB)