A Shared Subcircuit Lets LLMs Count Down Across Tasks
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
arXiv:2510.22752v2 Announce Type: replace-cross Abstract: In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual inf...
Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target.
arXiv:2605. 28854v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL).
The paper investigates how transformer models learn latent structure by training a small decoder-only transformer on three variants of the Alchemy benchmark. It finds that the model acquires different components of latent structure in discrete stages, with a notable asymmetry: it robustly composes fundamental transitions but struggles to decompose complex examples into atomic transitions. Layer‑specific causal interventions reveal plasticity windows where freezing layers delays or prevents stage completion, offering a detailed view of capability evolution during training.
arXiv:2609.23065v1 Announce Type: cross Abstract: Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through inp...
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
arXiv:2603.18908v5 Announce Type: replace Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
The paper investigates whether few-shot in-context learning (ICL) emerges similarly across different data modalities. Using a controlled cross-modality framework, the authors test the Convergent Emergence Hypothesis, which posits that tasks benefiting from ICL in one modality will also benefit in others. They find that paired-mapping ICL appears in six modalities—language, genome, integer sequences, time series, images, and proteins—outperforming baselines and showing correlated task effects in five of them, supporting the hypothesis in some but not all cases.
arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...
The paper explains why GPT‑style language models fail to transfer directly to symbolic music. It argues that success in language comes from tokenization that compresses data by creating a coordinate system where recurring patterns become predictable. For music, the authors propose that tokenization must build a predictively effective, relationally lossless coordinate system—defining Fact–Token and Token–State boundaries—to enable compression without sacrificing contextual freedom. Controlled experiments confirm that proper coordinate construction improves predictive compressibility, whereas mere sequence compaction does not.