Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
arXiv:2603. 13259v2 Announce Type: replace-cross Abstract: When a decoder-only transformer is forced to process matched correct and incorrect single-token continuations of a factual query, the two pathways through hidden-state space diverge in a specific way: displacement vectors from the query-only representation maintain approximately equal magnitude but rotate apart in direction.
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
arXiv:2607. 04926v1 Announce Type: cross Abstract: How does the way information reaches a transformer -- as symbolic tokens, a clean per-factor "oracle" code, or an entangled perceptual vector -- shape whether it binds that information compositionally?
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
arXiv:2605. 05686v3 Announce Type: replace Abstract: Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory, WM).
arXiv:2608. 03263v1 Announce Type: cross Abstract: We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data.
arXiv:2607. 12166v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms.
arXiv:2606. 07559v2 Announce Type: replace-cross Abstract: Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor.
arXiv:2608. 11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap.
arXiv:2606. 00926v1 Announce Type: new Abstract: Mechanistic studies of sequence models often treat layerwise state encodings as architectural traits: recurrent models concentrate readable state, attention-based models distribute it.
arXiv:2605. 18909v2 Announce Type: replace Abstract: Any system that models the world under finite representational capacity must compress; any compression entails a prior; and the prior is the system's bias.
arXiv:2606. 01060v1 Announce Type: cross Abstract: Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally.