arXiv:2609.08133v1 Announce Type: cross
Abstract: In nonconvex optimization problems arising in geometric machine learning, data augmentation is commonly used to promote invariance by averaging empir...
By Behrooz Tahmasebi, Melanie Weber
arXiv:2609.07031v1 Announce Type: new
Abstract: Learning with group invariances is central to many scientific and geometric learning problems, yet its computational foundations remain poorly understo...
By Ashkan Soleymani, Behrooz Tahmasebi, Patrick Jaillet, Stefanie Jegelka
The paper investigates sparse data augmentation for nonconvex optimization in geometric machine learning. It shows that using a small, fixed sample of transformations—obtained before optimization—allows gradient descent to achieve an ε‑stationary point of the fully augmented objective with ≤ O((log|G|+log(1/δ))/ε²) transformation queries. This is more efficient than both full augmentation and standard group‑SGD, which require O(1/ε⁴) queries.
The paper investigates self‑supervised pre‑training that uses multiple data augmentations of the same unlabeled sample. It shows that pooling these dependent augmentations together yields statistical estimation error bounds that are never worse than, and sometimes better than, partitioning the data into independent subsets. The analysis explains why using many augmentations is practically advantageous, especially when their correlations have mild effects or reduce estimation variance.
By Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe
arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic