The paper introduces two new tokenisation algorithms—BottomUpLL and TopDownComp—to systematically explore the 2x2 design space defined by optimisation objective (compression vs. log‑likelihood) and search procedure (bottom‑up merging vs. top‑down pruning). Experiments across model sizes, vocabularies, and domains show that the search procedure, rather than the objective, consistently yields lower bits‑per‑byte, while no clear pattern emerges on the BLiMP benchmark. These findings clarify how tokeniser design choices influence language‑model performance and provide guidance for constructing tokenisers more principledly.
By Ahmetcan Yavuz, Clara Meister, Tiago Pimentel
arXiv:2609.35869v1 Announce Type: new
Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only und...
By Yuhao Du, Shunian Chen
arXiv:2506. 15138v2 Announce Type: replace-cross Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost.
By Gyeongje Cho, Yeonkyoung So, Sangmin Lee, Jaejin Lee
The paper introduces Universal Byte-Level Encoding (UBE), a dual‑alphabet tokenizer that routes 3‑4‑byte UTF‑8 characters through UTF‑16 while keeping 1‑2‑byte characters on the UTF‑8 path. This design lowers the worst‑case token‑budget disparity for high‑premium scripts without increasing costs for efficient English spans, and it preserves standard BPE merges and exact decoding. In extensive Unicode audits and multilingual language‑model experiments, UBE matches or improves token‑count efficiency and context usability compared to traditional byte‑pair encoding.
By Hyunsik Kim, Youngmoon Jung
TokEval is a tokenizer evaluation suite that introduces metrics beyond traditional fertility and compression rate, capturing linguistically and structurally meaningful properties such as UTF‑8 character boundary integrity and digit place‑value alignment for mathematics. The authors validate these metrics by pretraining language models with varied tokenizers and measuring downstream performance on bits‑per‑byte and benchmarks covering linguistic understanding, mathematical reasoning, and code generation. Their results show that information‑theoretic metrics predict language modeling performance, while structure‑sensitive metrics correlate with task accuracy, suggesting TokEval can guide tokenizer selection more principledly.
By Clara Meister
arXiv:2511. 20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference.
By Dong Dong, Weijie Su