Hugging Face Blog

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

arXiv AI
Jun 9

Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications

arXiv:2603. 22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the functional contribution of each component type remains insufficiently characterized.

By Hector Borobia, Elies Segu\'i-Mas, Guillermina Tormo-Carb\'o
arXiv Computation and Language
Sep 7

What Attention Recalls and Recurrence Controls in Hybrid Language Models

Hybrid language models combine attention with a fixed-size recurrent state, yet the distinct roles of each component are not well understood. The authors introduce two cache-level interventions—split-prefill and state-swap—to isolate the contributions of the KV cache (attention) and the recurrent state. Experiments on Qwen3.5 and Falcon-H1 show that exact retrieval depends almost entirely on attention, while output language and persona rely mainly on recurrence, with the state-swap intervention confirming that answers derive from the KV side and language from the recurrent side.

By Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov
arXiv AI
Sep 25

An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model

The paper reports a single‑seed ablation study of the TALH language model, which combines a Multi‑head Latent Attention (MLA) branch with a custom recurrent state‑space model (SSM). Five variants ranging from 117 M to 217 M active parameters per token were trained on a FineWeb sample, and the results show that removing the SSM branch causes the largest drop in validation perplexity (315) compared to removing MLA (239). A dense‑FFN hybrid achieved a perplexity of 231, outperforming the tested top‑2 ternary‑MoE hybrid (240) while using 3.87 GB less peak training memory, and MLA‑only exhibited the flattest time‑to‑first‑token curve on an Apple M3, though the dense Transformer was faster overall.

By Christos Koutsiaris
arXiv Computation and Language
Sep 2

Toppling the Hierarchy in Byte-level Language Modeling

The paper investigates why current byte‑level language models, which use a hierarchical structure that down‑samples to words and then upsamples back to bytes, struggle with precise character manipulation. Experiments show that pure byte‑level models outperform hierarchical variants on character‑level tasks, and that byte‑level attention is the key component driving this advantage. The study explains the trade‑off between computational efficiency and fine‑grained character understanding in hierarchical byte models.

By Lukas Edman, Alexander Fraser
arXiv Computation and Language
Sep 25

Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.

By Zhenyan Lu, He Wang, Xiaohui Huang