Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing
arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
arXiv:2604. 03444v4 Announce Type: replace Abstract: Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attention.
arXiv:2606. 27449v1 Announce Type: new Abstract: Multi-head attention conventionally partitions the hidden dimension equally across all heads at every layer, enforcing an identical representational subspace dimension (dh = dmodel/h) throughout the models depth.
arXiv:2608. 19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets.
arXiv:2601. 11667v2 Announce Type: replace-cross Abstract: Transformer architectures deliver state-of-the-art accuracy via dense full-attention, but their quadratic time and memory complexity with respect to sequence length limits practical deployment.
arXiv:2606. 15378v1 Announce Type: cross Abstract: Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention modules, such as sliding-window attention (SWA) and recurrent sequence mixers.
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.
arXiv:2602. 21196v2 Announce Type: replace Abstract: Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism.
The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.
The paper introduces Post-Boundary Bridge (PBB), a hybrid transformer architecture that keeps causal attention within blocks while adding direct connections across block boundaries. PBB focuses on within-block modeling and nearby exchange, delegating long-range communication to full-attention layers. Experiments on dense and mixture-of-experts models ranging from 205 million to 2.07 billion parameters show that PBB hybrids maintain near-full perplexity, competitive downstream performance, and improved source retrieval, while Flash-PBB achieves 1.82× faster decoding with half the local key‑value cache compared to Flash‑SWA.
arXiv:2605.26797v2 Announce Type: replace Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden...
arXiv:2502.09245v3 Announce Type: replace Abstract: In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standar...