arXiv Statistics ML

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

arXiv Machine Learning
Sep 25

Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

Transformers can learn broad families of tasks during pretraining and adapt to unseen tasks from a short prompt, but a rigorous understanding of this capability is limited. This paper studies how shared cross‑task structure influences the sample complexity of in‑context learning (ICL) by characterizing task‑space complexity through covering numbers, yielding a set of anchor functions that localize unseen tasks and predict responses. The authors construct a Transformer with Softmax attention to approximate this procedure and derive an error bound that separates the effects of pretraining tasks and prompt length, showing that once enough tasks are available the dependence on prompt length becomes dimension‑free.

By Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao
arXiv Machine Learning
Sep 7

Nested Inductive Bias Framework for SPD Manifold Learning

The paper introduces a Nested Inductive Bias framework that uses a two‑stage diffeomorphic composition to pull back non‑Euclidean target geometries onto symmetric positive definite (SPD) manifolds. This approach allows the construction of curvature‑aligned Riemannian classifiers that respect both matrix constraints and the intrinsic relational geometry of data. Empirical results on kinematic, signal processing, and synthetic benchmarks show that class separability degrades when metric curvature does not match the data distribution, and the authors also propose the Rational Conformal Metric (RCM) for robust vectorized architectures.

By Tushar Das
arXiv Machine Learning
Jun 9

Generalization in Nonlinear Least Squares via Learned Feature Geometry

arXiv:2606. 08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-dependent effective dimension that reflects the geometry of the gradient model at the trained parameters, through the empirical Jacobian Gram matrix and a residual--curvature term.

By Ayub Kharel, Ilja Kuzborski, Patrick Rebeschini, Yasin Abbasi-Yadkori