The State-Prediction Separation Hypothesis
arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.
The paper introduces Sparse Token Routing in Efficient Transformers, evaluating a two-stream Transformer (SEWN) that routes tokens through either lightweight or full-capacity processing via a learned gate. Experiments show that routing causes negligible accuracy change compared to parameter-matched baselines, and that the effectiveness of the gate’s token-importance signal depends on its learning method. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without altering task accuracy.
arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.
arXiv:2607. 22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules.
arXiv:2609.27233v1 Announce Type: new Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.
arXiv:2609.15131v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most method...
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
T-LoopFormer introduces token-level elastic-depth looped transformers that allow each token to decide its own number of loop iterations based on its hidden state, improving token generation accuracy. It also adds a recursion-wise key‑value cache so tokens at different depths only attend to their corresponding cached states, speeding up autoregressive decoding. Experiments demonstrate strong performance on language modeling and zero‑shot reasoning, achieving the lowest decoding latency among comparable models.
arXiv:2604. 11912v2 Announce Type: replace-cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks.
arXiv:2502.09245v3 Announce Type: replace Abstract: In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standar...
arXiv:2605.21333v3 Announce Type: replace-cross Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combinati...
arXiv:2608.29291v1 Announce Type: new Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generati...
TopK-Guided is a training‑free method that improves activation sparsity for large language model inference by combining token‑level sparsity adaptation with block‑level budget allocation that accounts for block sensitivity. It addresses limitations of existing methods like TEAL, which adapts sparsity per token but lacks tight control, and WINA, which enforces a fixed sparsity across all tokens and blocks. Experiments on Llama‑2 and Llama‑3 show that TopK‑Guided consistently yields better perplexity and downstream accuracy while maintaining similar compute costs to WINA, especially at high sparsity levels.