arXiv Machine Learning

Data Augmentation: A Fourier Analysis Perspective

arXiv:2606. 24418v1 Announce Type: new Abstract: Data augmentation is a simple and model-agnostic approach for exploiting known invariances in learning problems.

Hugging Face Trending Papers
Sep 8

Sparse Data Augmentation for Optimization with Provable Guarantees

The paper investigates sparse data augmentation for nonconvex optimization in geometric machine learning. It shows that using a small, fixed sample of transformations—obtained before optimization—allows gradient descent to achieve an ε‑stationary point of the fully augmented objective with ≤ O((log|G|+log(1/δ))/ε²) transformation queries. This is more efficient than both full augmentation and standard group‑SGD, which require O(1/ε⁴) queries.

arXiv Machine Learning
Sep 7

An Analysis of Self-supervised Pre-training with Dependent Samples

The paper investigates self‑supervised pre‑training that uses multiple data augmentations of the same unlabeled sample. It shows that pooling these dependent augmentations together yields statistical estimation error bounds that are never worse than, and sometimes better than, partitioning the data into independent subsets. The analysis explains why using many augmentations is practically advantageous, especially when their correlations have mild effects or reduce estimation variance.

By Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe
Hugging Face Trending Papers
Jun 8

Data augmented bootstrap: Unifying confidence interval construction by approximate invariance

We propose the data augmented bootstrap (DAB), a framework for constructing confidence intervals from approximately invariant transformations of the data. As special cases, DAB recovers popular methods that rely on exact group symmetries, such as conformal prediction, wild bootstrap for Maximum Mean Discrepancy U-statistics and the recently proposed SymmPI.

arXiv AI
Jun 30

Representation Learning for Equivariant Inference with Guarantees

arXiv:2505. 19809v3 Announce Type: replace-cross Abstract: In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries rooted in physics or geometry can dramatically improve generalization and sample efficiency.

By Daniel Ordo\~nez-Apraez, Vladimir Kosti\'c, Alek Fr\"ohlich, Vivien Brandt, Karim Lounici, Massimiliano Pontil
arXiv Machine Learning
Aug 28

Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning

The paper investigates fundamental limits of algorithmic principles in multiclass learning, specifically proper learning and regularization. It shows that learning cannot always be reduced to proper learning even with an enlarged hypothesis class, that proper learners may need a sublinear number of errors that can be arbitrarily large, and that regularization (SRM or local) is not universally sufficient. The authors also provide a positive theory giving sufficient conditions for SRM learnability and a characterization via integrability of revealed preferences.

By Julian Asilis, Shaddin Dughmi, Vatsal Sharan, Alec Sun, Shang-Hua Teng, Chang Wang
arXiv Machine Learning
Jun 2

Symmetries in PAC-Bayesian Learning

arXiv:2510. 17303v2 Announce Type: replace Abstract: Symmetries are known to improve the empirical performance of machine learning models, yet theoretical guarantees explaining these gains remain limited.

By Armin Beck, Peter Ochs