Welcome Falcon Mamba: The first strong attention-free 7B model
Related stories
Codestral Mamba
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
InfoMamba is an attention‑free hybrid model that combines a minimal‑bandwidth global interface with a selective recurrent stream. The architecture replaces token‑level self‑attention with a concept bottleneck linear filtering layer and integrates it via an information‑maximizing fusion (IMF) that injects global context into the state‑space dynamics. Experiments across classification, dense prediction, and non‑vision tasks show that InfoMamba outperforms strong Transformer and SSM baselines while maintaining near‑linear scaling and competitive accuracy‑efficiency trade‑offs.
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.
Zamba2-VL Technical Report
arXiv:2606. 00390v1 Announce Type: cross Abstract: We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with a small number of shared transformer blocks.
Understanding BigBird's Block Sparse Attention
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
arXiv:2608. 02347v2 Announce Type: replace Abstract: Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling.
Falcon 2: An 11B parameter pretrained language model and VLM, trained on over 5000B tokens and 11 languages
A Visual Guide to Attention Variants in Modern LLMs
From MHA and GQA to MLA, sparse attention, and hybrid architectures
An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals
arXiv:2607. 11796v1 Announce Type: new Abstract: Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism.
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink
arXiv:2606.00930v2 Announce Type: replace-cross Abstract: Mechanistic interpretability routinely reads a probe and labels its top-activating units as the circuit executing the computation. We test th...
