Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing
arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
arXiv:2606. 24650v1 Announce Type: cross Abstract: We present Harmonic, a hierarchical state space model (SSM) for language modeling.
arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
The paper reports a single‑seed ablation study of the TALH language model, which combines a Multi‑head Latent Attention (MLA) branch with a custom recurrent state‑space model (SSM). Five variants ranging from 117 M to 217 M active parameters per token were trained on a FineWeb sample, and the results show that removing the SSM branch causes the largest drop in validation perplexity (315) compared to removing MLA (239). A dense‑FFN hybrid achieved a perplexity of 231, outperforming the tested top‑2 ternary‑MoE hybrid (240) while using 3.87 GB less peak training memory, and MLA‑only exhibited the flattest time‑to‑first‑token curve on an Apple M3, though the dense Transformer was faster overall.
The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.
arXiv:2606. 26290v1 Announce Type: cross Abstract: While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accumulation remains under-explored.
arXiv:2604. 00004v2 Announce Type: replace-cross Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT).
arXiv:2605.21333v3 Announce Type: replace-cross Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combinati...
arXiv:2606. 28876v1 Announce Type: cross Abstract: Long-context language models often conflate two different goals: compressing history into an efficient state, and maintaining reliable long-term memory.
arXiv:2607. 19368v1 Announce Type: new Abstract: Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length.
arXiv:2606. 18694v1 Announce Type: new Abstract: A network of oscillators that synchronizes perfectly computes nothing further, so an attention architecture built from synchronization must locate its computation in structured departures from agreement.
arXiv:2605. 21333v2 Announce Type: replace-cross Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combination that has produced a persistent quality gap relative to dense Transformers.
arXiv:2607. 14427v1 Announce Type: new Abstract: A depth-recurrent transformer applies a weight-tied core a variable number of times, and prior work has shown that training with a randomized recursion count yields one checkpoint usable across a range of inference depths.
arXiv:2607. 01394v1 Announce Type: new Abstract: We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family including GPT, LLaMA, Mistral, or Falcon.