arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2609.22125v1 Announce Type: new
Abstract: Standard tokenizers used in large language models produce malformed text when applied to Brahmic scripts. They are a family of abugidas, writing system...
By Sai Hemanth Kapila, Rakshika Bagavathy
arXiv:2608. 00523v2 Announce Type: replace-cross Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon.
By Mohamed El Idrissi
This article tackles an important phenomenon in the syntax of Yemeni Ibbi Arabic (YIA), viz. , wh-agreement, a phenomenon common to several languages including Greek, Indonesian, Lubukusu, Irish, etc.
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
By Albin Th\"orn Cleland, Eric Cullhed
TatBLiMP is the first benchmark of linguistic minimal pairs for the Tatar language, covering 16 morphosyntactic phenomena across 1,248 sentence pairs that differ by a single morpheme. Each pair contains one grammatical and one ungrammatical sentence, with the ungrammatical version generated by a deterministic perturbation and ratified by a native speaker. The benchmark evaluates models by comparing their assigned probabilities, allowing assessment without text generation or parsing, and tracks performance across from-scratch, cross‑lingual, and multilingual large language models.
By Ilshat Saetov, Dmitry Gaynullin
The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.
By Jungyeul Park, KyungTae Lim, Zihao Huang, Eunkyul Leah Jo, Yige Chen, Chulwoo Park
The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
Language models trained on tokenized text still reliably produce morphemes whose form depends on phonology, but it was unclear whether this relies on memorization or rule-like generalization. The study shows that for the English indefinite article a/an, the phonological condition is encoded along a single linear direction in trigger-token embeddings, causally drives article selection in token-level wug tests, and is used by the model to forecast the upcoming trigger token’s phonological feature for article choice. The authors further investigate whether this rule-like generalization extends to allomorph selection in other languages and to explicit phonological judgment, offering a mechanistic account that separates generation-time ability from metalinguistic judgments.
By Sangwoo Kim, Sangah Lee
arXiv:2607. 18961v1 Announce Type: new Abstract: Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix.
By Remo Pareschi
The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.
The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.
By Giuseppe Samo, Vivi Nastase, Paola Merlo