arXiv AI

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

arXiv:2607. 15819v1 Announce Type: cross Abstract: In-context learning is a remarkable property of transformers and has recently received a lot of interest.

arXiv Machine Learning
Aug 27

Cubit: Token Mixer with Kernel Ridge Regression

The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.

By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu
arXiv Machine Learning
Sep 25

Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

Transformers can learn broad families of tasks during pretraining and adapt to unseen tasks from a short prompt, but a rigorous understanding of this capability is limited. This paper studies how shared cross‑task structure influences the sample complexity of in‑context learning (ICL) by characterizing task‑space complexity through covering numbers, yielding a set of anchor functions that localize unseen tasks and predict responses. The authors construct a Transformer with Softmax attention to approximate this procedure and derive an error bound that separates the effects of pretraining tasks and prompt length, showing that once enough tasks are available the dependence on prompt length becomes dimension‑free.

By Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao
arXiv Machine Learning
Sep 11

Test time training enhances in-context learning of nonlinear functions

The paper studies how test‑time training (TTT) improves in‑context learning (ICL) for nonlinear models, focusing on single‑index models where features lie in a hidden low‑dimensional subspace. By applying TTT to single‑layer transformers trained with gradient‑based methods, the authors derive an upper bound on prediction risk and show that TTT allows the model to adapt to both feature vectors and link functions that vary across tasks—something ICL alone struggles to achieve. They also provide a convergence rate indicating that predictive error can approach the noise level as context size and network width increase.

By Kento Kuwataka, Taiji Suzuki
arXiv Machine Learning
Sep 10

Conditioned Initialization for Attention

arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...

By Hemanth Saratchandran, Simon Lucey