arXiv:2511. 02496v2 Announce Type: replace Abstract: We study latent geometry as an explicit component of representation quality in data-scarce learning.
By Ronald Katende
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
arXiv:2606. 00512v1 Announce Type: new Abstract: In many modern machine learning pipelines, abundant pretrained representations serve as noisy proxy covariates, while task-specific labels remain scarce.
By Kwangho Kim, Jisu Kim
Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases. However, existing theoretical frameworks for pre-training do not fully explain this phenomenon.
arXiv:2609.31458v1 Announce Type: new
Abstract: Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large languag...
By Jaehee Seo, Jisu Kim
arXiv:2606. 24418v1 Announce Type: new Abstract: Data augmentation is a simple and model-agnostic approach for exploiting known invariances in learning problems.
By Behrooz Tahmasebi, Melanie Weber, Stefanie Jegelka
arXiv:2607. 07513v1 Announce Type: new Abstract: Self-supervised learning matches supervised accuracy from a fraction of the labels, but the labeled-sample efficiency behind this has lacked a theoretical explanation.
By Adam M. Oberman
arXiv:2510. 01163v2 Announce Type: replace Abstract: The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples.
By Wa\"iss Azizian, Ali Hasan
arXiv:2606. 02008v1 Announce Type: cross Abstract: Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases.
By Kazuto Fukuchi, Ryuichiro Hataya, Kota Matsui
arXiv:2606. 31208v1 Announce Type: new Abstract: Large tabular models (LTMs), i.
By Francesco Capano, Jonas B\"ohler
arXiv:2608. 06250v1 Announce Type: cross Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not.
By Alex Buna, Shirley Xiaoqi Liu, Patrick Rebeschini
arXiv:2605. 27991v2 Announce Type: replace-cross Abstract: Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early-stopping rules.
By Minhao Yao, Ruoyu Wang, Xihong Lin, Lin Liu, Zhonghua Liu