arXiv AI By Tolga \c{S}akar

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

Read the original on arXiv AI →

arXiv:2606. 18717v1 Announce Type: cross Abstract: Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

MoganBert-TR is a 149‑million‑parameter Turkish encoder foundation model trained from scratch on a language‑specific corpus using a two‑stage CLM‑to‑MLM curriculum. The model, along with its embedding variant MoganBert‑Embed, achieves state‑of‑the‑art results on Turkish benchmarks such as TrGLUE and TabiBench, outperforming existing Turkish BERT models. Its tokenizer, comprising 50,048 tokens, also surpasses other Turkish tokenizers in compression and fertility metrics.

By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
arXiv Computation and Language
Sep 24

Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings

The paper argues that meaning identity—whether two sentences convey the same idea after wording changes—is not encoded in the geometry of independently produced sentence embeddings. Experiments on frozen off‑the‑shelf encoders and language models show that identity can only be reliably computed when both sentences are processed together in a single forward pass, yielding high accuracy (0.90–0.96) on PAWS‑X, whereas independent embeddings or simple fusion methods perform near chance. Even advanced bi‑encoder fine‑tuning improves performance on PAWS but fails to generalize to other similarity tasks, underscoring that identity is a cheap computed operator rather than a property of individual sentence vectors.

By Jiaqi Deng
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv Machine Learning
Aug 28

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.

By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv AI
Aug 11

Length-MAX Tokenizer for Language Models

arXiv:2511. 20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference.

By Dong Dong, Weijie Su