arXiv:2608. 00523v2 Announce Type: replace-cross Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon.
By Mohamed El Idrissi
arXiv:2603.03510v3 Announce Type: replace-cross
Abstract: This study investigates the diverse characteristics of nouns, focusing on both semantic (e.g., countable/uncountable) and morphosyntactic (e....
By Mohamed El Idrissi
arXiv:2606. 24172v1 Announce Type: cross Abstract: More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped.
By Ritwik Banerjee, Lav R. Varshney
TatBLiMP is the first benchmark of linguistic minimal pairs for the Tatar language, covering 16 morphosyntactic phenomena across 1,248 sentence pairs that differ by a single morpheme. Each pair contains one grammatical and one ungrammatical sentence, with the ungrammatical version generated by a deterministic perturbation and ratified by a native speaker. The benchmark evaluates models by comparing their assigned probabilities, allowing assessment without text generation or parsing, and tracks performance across from-scratch, cross‑lingual, and multilingual large language models.
By Ilshat Saetov, Dmitry Gaynullin
The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.
By Elzbieta Dawidek (University of Lower Silesia DSW Ideis)
The article investigates Korean predicate morphology, showing that a sequence of canonical morphemes and grammatical labels does not uniquely determine the surface form for certain predicates. It demonstrates that identical or nearly identical stem-ending configurations can produce different outputs depending on lexical identity and realization class membership. The study frames this as homonymy with inflectional divergence, highlighting that lexical meaning, subcategorization, and semantic role structure are essential for determining the correct surface realization.
By Wonjun Oh, KyungTae Lim, Jungyeul Park
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2606. 15144v1 Announce Type: cross Abstract: Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure that underlies word formation.
By Jann Railey Montalan, David Demitri Africa, Jimson Paulo Layacan, Richell Isaiah Flores, Ivan Yuri De Leon, Lance Calvin Gamboa
The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.
By Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha
Language models trained on tokenized text still reliably produce morphemes whose form depends on phonology, but it was unclear whether this relies on memorization or rule-like generalization. The study shows that for the English indefinite article a/an, the phonological condition is encoded along a single linear direction in trigger-token embeddings, causally drives article selection in token-level wug tests, and is used by the model to forecast the upcoming trigger token’s phonological feature for article choice. The authors further investigate whether this rule-like generalization extends to allomorph selection in other languages and to explicit phonological judgment, offering a mechanistic account that separates generation-time ability from metalinguistic judgments.
By Sangwoo Kim, Sangah Lee
The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.
By Jungyeul Park, KyungTae Lim, Zihao Huang, Eunkyul Leah Jo, Yige Chen, Chulwoo Park