← Back to all news
Hugging Face Blog December 18, 2025

Tokenization in Transformers v5: Simpler, Clearer, and More Modular

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms
  • nlp

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Feb 14, 2020

How to train a new language model from scratch using Transformers and Tokenizers

llms
More like this →
arXiv AI
Jul 9

Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

arXiv:2607. 07066v1 Announce Type: cross Abstract: Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition.

By Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan, Hemkesh Bandi, Jarod Alper
llmsrag
More like this →
arXiv AI
Jun 19

Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers

arXiv:2606. 20076v1 Announce Type: cross Abstract: Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio.

By Dong Hoon Lee, Seunghoon Hong
llmsdiffusionnlpsafety
More like this →
arXiv Machine Learning
Jul 7

Equivalence of Context and Parameter Updates in Modern Transformer Blocks

arXiv:2511. 17864v3 Announce Type: replace Abstract: Recent research has established that the impact of context in a vanilla transformer can be represented implicitly by forming a token-dependent, rank-1 patch to its MLP weights.

By Adrian Goldwaser, Michael Munn, Javier Gonzalvo, Benoit Dherin
llms
More like this →
arXiv AI
Aug 14

On the Expressive Power of Transformers

arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.

By Phokion Kolaitis, Rik Sengupta
llms
More like this →
arXiv Machine Learning
Jun 9

Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions

arXiv:2606. 08768v1 Announce Type: new Abstract: Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings.

By Blanka K\"over, Alexandra Butoi, Anej Svete, Michael Hahn, Ryan Cotterell
llmssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e