Welcome Falcon Mamba: The first strong attention-free 7B model
Related stories
Codestral Mamba
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.
Zamba2-VL Technical Report
arXiv:2606. 00390v1 Announce Type: cross Abstract: We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with a small number of shared transformer blocks.
Understanding BigBird's Block Sparse Attention
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
arXiv:2608. 02347v2 Announce Type: replace Abstract: Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling.
Falcon 2: An 11B parameter pretrained language model and VLM, trained on over 5000B tokens and 11 languages
A Visual Guide to Attention Variants in Modern LLMs
From MHA and GQA to MLA, sparse attention, and hybrid architectures
An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals
arXiv:2607. 11796v1 Announce Type: new Abstract: Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism.
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data
arXiv:2605. 22775v2 Announce Type: replace-cross Abstract: Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handling frequent data missingness from blinks and tracking failures, and efficiently modeling long-range temporal dependencies.
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
arXiv:2606. 00930v1 Announce Type: cross Abstract: Mechanistic interpretability often assumes that probes identifying a representational signature also identify the circuit executing the corresponding computation.
