Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Related stories
Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications
arXiv:2603. 22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the functional contribution of each component type remains insufficiently characterized.
Falcon 2: An 11B parameter pretrained language model and VLM, trained on over 5000B tokens and 11 languages
Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
Falcon-Arabic: A Breakthrough in Arabic Language Models
Large Language Models: A New Moore's Law?
Understanding the capabilities, limitations, and societal impact of large language models
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs.
First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers
The paper introduces Language Identity Head Ablation (LIHA), a causal method that zeroes individual attention heads in transformer models to measure language switch rates across multilingual prompts. Applying LIHA to GPT‑2 reveals a small set of first‑token broadcaster heads—most notably L6H1—that persistently attend to the initial prompt token and propagate language signals throughout generation, with compensatory head recruitment occurring hierarchically in higher layers. A controlled comparison between Qwen2.5‑1.5B‑Base and Qwen2.5‑1.5B‑Instruct shows that instruction tuning concentrates language‑identity influence in early layers, while experiments with Chinese and Russian confirm script‑specific first‑token broadcasting at layer 0.
Red-Teaming Large Language Models
Durable Evaluation Framework: Adversarial Arbitration for Sycophancy Reduction in Large Language Models
arXiv:2606. 07532v2 Announce Type: replace-cross Abstract: RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process.
Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian
arXiv:2402. 18121v2 Announce Type: replace-cross Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language.