arXiv Machine Learning

In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning

arXiv:2510. 10981v3 Announce Type: replace-cross Abstract: This paper develops a finite-sample statistical theory for in-context learning (ICL), analyzed within a meta-learning framework that accommodates mixtures of diverse task types.

arXiv Machine Learning
Sep 25

Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

Transformers can learn broad families of tasks during pretraining and adapt to unseen tasks from a short prompt, but a rigorous understanding of this capability is limited. This paper studies how shared cross‑task structure influences the sample complexity of in‑context learning (ICL) by characterizing task‑space complexity through covering numbers, yielding a set of anchor functions that localize unseen tasks and predict responses. The authors construct a Transformer with Softmax attention to approximate this procedure and derive an error bound that separates the effects of pretraining tasks and prompt length, showing that once enough tasks are available the dependence on prompt length becomes dimension‑free.

By Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao
arXiv Machine Learning
Jul 7

A Unified Framework for In-Context Learning with Causal and Masked Language Models

arXiv:2607. 04081v1 Announce Type: new Abstract: In-context learning (ICL) has emerged as a central capability of pretrained language models, yet its theoretical analysis has focused primarily on causal language models trained by left-to-right autoregressive prediction, such as GPT-style models.

By Chenrui Liu, Chuanlong Xie, Falong Tan, Yicheng Zeng, Lixing Zhu