Are In-Context Images Worth 10 Dimensions?
arXiv:2609.37659v1 Announce Type: cross Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.
arXiv:2609.37659v1 Announce Type: cross Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction c...
Transformers can learn broad families of tasks during pretraining and adapt to unseen tasks from a short prompt, but a rigorous understanding of this capability is limited. This paper studies how shared cross‑task structure influences the sample complexity of in‑context learning (ICL) by characterizing task‑space complexity through covering numbers, yielding a set of anchor functions that localize unseen tasks and predict responses. The authors construct a Transformer with Softmax attention to approximate this procedure and derive an error bound that separates the effects of pretraining tasks and prompt length, showing that once enough tasks are available the dependence on prompt length becomes dimension‑free.
arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
arXiv:2607. 10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge.
arXiv:2607. 00479v1 Announce Type: new Abstract: Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
arXiv:2606. 27449v1 Announce Type: new Abstract: Multi-head attention conventionally partitions the hidden dimension equally across all heads at every layer, enforcing an identical representational subspace dimension (dh = dmodel/h) throughout the models depth.
arXiv:2609.08981v1 Announce Type: cross Abstract: A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing...
arXiv:2605. 28854v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL).
Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.