arXiv Machine Learning By Hyunsik Kim, Youngmoon Jung

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

Read the original on arXiv Machine Learning →

The paper introduces Universal Byte-Level Encoding (UBE), a dual‑alphabet tokenizer that routes 3‑4‑byte UTF‑8 characters through UTF‑16 while keeping 1‑2‑byte characters on the UTF‑8 path. This design lowers the worst‑case token‑budget disparity for high‑premium scripts without increasing costs for efficient English spans, and it preserves standard BPE merges and exact decoding. In extensive Unicode audits and multilingual language‑model experiments, UBE matches or improves token‑count efficiency and context usability compared to traditional byte‑pair encoding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

Objective vs. Search: Decomposing What Makes a Good Tokeniser

The paper introduces two new tokenisation algorithms—BottomUpLL and TopDownComp—to systematically explore the 2x2 design space defined by optimisation objective (compression vs. log‑likelihood) and search procedure (bottom‑up merging vs. top‑down pruning). Experiments across model sizes, vocabularies, and domains show that the search procedure, rather than the objective, consistently yields lower bits‑per‑byte, while no clear pattern emerges on the BLiMP benchmark. These findings clarify how tokeniser design choices influence language‑model performance and provide guidance for constructing tokenisers more principledly.

By Ahmetcan Yavuz, Clara Meister, Tiago Pimentel