arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
By Kuzey Torlak, H\"useyin Arda Arslan, An{\i}l Dervi\c{s}o\u{g}lu, Beyza Nur Deniz, Onur Boyar
The paper reports a single‑seed ablation study of the TALH language model, which combines a Multi‑head Latent Attention (MLA) branch with a custom recurrent state‑space model (SSM). Five variants ranging from 117 M to 217 M active parameters per token were trained on a FineWeb sample, and the results show that removing the SSM branch causes the largest drop in validation perplexity (315) compared to removing MLA (239). A dense‑FFN hybrid achieved a perplexity of 231, outperforming the tested top‑2 ternary‑MoE hybrid (240) while using 3.87 GB less peak training memory, and MLA‑only exhibited the flattest time‑to‑first‑token curve on an Apple M3, though the dense Transformer was faster overall.
By Christos Koutsiaris
The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.
By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko
arXiv:2606. 26290v1 Announce Type: cross Abstract: While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accumulation remains under-explored.
By Omanshu Thapliyal
arXiv:2604. 00004v2 Announce Type: replace-cross Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT).
By Ning Yang, Hengyu Zhong, Wentao Wang, Baoliang Tian, Haijun Zhang, Jun Wang
arXiv:2605.21333v3 Announce Type: replace-cross
Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combinati...
By Ting Liu