← Back to all news
Hugging Face Blog December 18, 2025

Tokenization in Transformers v5: Simpler, Clearer, and More Modular

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms
  • nlp

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Feb 14, 2020

How to train a new language model from scratch using Transformers and Tokenizers

llms
More like this →
Hugging Face Blog
Sep 21

tokenizers v1: encode, decode and scaling, measured

More like this →
Hugging Face Trending Papers
Sep 23

Distilling Sequential Computation in Transformer Language Models

Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...

llmsragnlpefficiency
More like this →
arXiv AI
Jul 9

Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

arXiv:2607. 07066v1 Announce Type: cross Abstract: Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition.

By Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan, Hemkesh Bandi, Jarod Alper
llmsrag
More like this →
arXiv AI
Jun 19

Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers

arXiv:2606. 20076v1 Announce Type: cross Abstract: Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio.

By Dong Hoon Lee, Seunghoon Hong
llmsdiffusionnlpsafety
More like this →
arXiv Computation and Language
Sep 24

Distilling Sequential Computation in Transformer Language Models

arXiv:2609.27233v1 Announce Type: new Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...

By Zixuan Lan, Jessica Yang, Yanhong Li, Karen Livescu, Jiawei Zhou
llmsragnlpefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea