A Shared Subcircuit Lets LLMs Count Down Across Tasks
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target.
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains.
The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.
The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
The paper investigates how LLaMA 3.1‑8B models numerical sequence patterns, focusing on time‑series prediction. By designing a task that requires detecting structural cues—specifically first differences in a sequence—the authors show that the model performs well and internally computes and stores these differences. Probing and activation‑patching experiments reveal that LLaMA retrieves and applies the first‑difference via an induction‑like circuit, marking one of the first demonstrations of concept induction in large language models.
arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.
arXiv:2606. 15521v1 Announce Type: cross Abstract: Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string.
arXiv:2609.34187v2 Announce Type: replace-cross Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
The paper introduces DRY, a sampling-time logit adjustment that penalizes token generation only when it would extend the current suffix into an exact repetition of an earlier span, thereby preventing verbatim loops in large language model outputs. Experiments across models ranging from 1.5B to 120B parameters and various prompt families show that DRY cuts suffix-extension rates by 47% and improves lexical diversity, while preserving benchmark performance. The method has been adopted by popular open-source LLM inference frameworks, indicating its practical relevance.
arXiv:2510.22752v2 Announce Type: replace-cross Abstract: In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual inf...
arXiv:2609.27233v1 Announce Type: new Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...