Falcon 2: An 11B parameter pretrained language model and VLM, trained on over 5000B tokens and 11 languages
Related stories
Length-MAX Tokenizer for Language Models
arXiv:2511. 20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference.
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Matryoshka Language Model Suites
arXiv:2608. 09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently.
SEA-LION-v4.8: A Technical Report
The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.
IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.
The Life of a Token: from Words to Bits on the Wire
The article "The Life of a Token: from Words to Bits on the Wire" explores how large language models convert text into network traffic during training. It traces the transformation from words to tokens, then to vectors, and finally to binary streams that traverse high‑performance computing systems. Using Dante’s Divine Comedy as a case study, the tutorial examines how tokenization, embeddings, and parallelization affect the volume, structure, and timing of data exchanged across the network, and provides analytical traffic models and numerical examples to clarify the communication demands of LLM training.
TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
arXiv:2608.30567v1 Announce Type: new Abstract: We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-co...
Prototype Language Models
arXiv:2607. 00510v1 Announce Type: new Abstract: Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc.
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
arXiv:2607. 08974v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the backbone so extensively that it becomes difficult to isolate what the VLM contributes to control.
TokEval: A Tokenizer Evaluation Suite
TokEval is a tokenizer evaluation suite that introduces metrics beyond traditional fertility and compression rate, capturing linguistically and structurally meaningful properties such as UTF‑8 character boundary integrity and digit place‑value alignment for mathematics. The authors validate these metrics by pretraining language models with varied tokenizers and measuring downstream performance on bits‑per‑byte and benchmarks covering linguistic understanding, mathematical reasoning, and code generation. Their results show that information‑theoretic metrics predict language modeling performance, while structure‑sensitive metrics correlate with task accuracy, suggesting TokEval can guide tokenizer selection more principledly.
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
arXiv:2602. 23881v2 Announce Type: replace Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate tokens that are then verified in parallel by the target model.