arXiv Machine Learning

Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

arXiv:2607. 09250v1 Announce Type: cross Abstract: The impact of a given training point on a statistical model is classically measured through its leave-one-out influence, which quantifies the effect of its removal from the training set on the model accuracy.

arXiv Machine Learning
Jun 15

A Complexity Measure for Active Learning in Multi-group Mean Estimation

arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.

By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
arXiv Machine Learning
Sep 21

Sparse Priors for Efficient Distribution Learning

arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.

By Saumya Goyal, Barnab\'as P\'oczos
arXiv Machine Learning
Jun 3

Testing Most Influential Sets

arXiv:2510. 20372v4 Announce Type: replace-cross Abstract: Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings.

By Lucas D. Konrad, Nikolas Kuschnig
arXiv Machine Learning
2d ago

How Many Samples Are Enough for Learning Across Domains?

The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.

By Hong Zheng
arXiv Machine Learning
Jul 14

Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds

arXiv:2607. 10592v1 Announce Type: new Abstract: Many geometric statistics and manifold learning pipelines routinely produce observations -- such as tangent vectors or local frames -- whose natural home is a varying family of fibers attached to different points of a base manifold, rather than a single shared vector space.

By Swagatam Das, Vaclav Snasel
arXiv Machine Learning
Jun 16

Active Learning with Low-Rank Structure for Data Selection

arXiv:2606. 16045v1 Announce Type: new Abstract: In the data selection problem, the objective is to choose a small, representative subset of data that can be used to efficiently train a machine learning model.

By Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan, David P. Woodruff, Samson Zhou
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto