Hugging Face Trending Papers

The syntax of wh-agreement in Yemeni Ibbi Arabic

This article tackles an important phenomenon in the syntax of Yemeni Ibbi Arabic (YIA), viz. , wh-agreement, a phenomenon common to several languages including Greek, Indonesian, Lubukusu, Irish, etc.

arXiv Computation and Language
Sep 21

TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

TatBLiMP is the first benchmark of linguistic minimal pairs for the Tatar language, covering 16 morphosyntactic phenomena across 1,248 sentence pairs that differ by a single morpheme. Each pair contains one grammatical and one ungrammatical sentence, with the ungrammatical version generated by a deterministic perturbation and ratified by a native speaker. The benchmark evaluates models by comparing their assigned probabilities, allowing assessment without text generation or parsing, and tracks performance across from-scratch, cross‑lingual, and multilingual large language models.

By Ilshat Saetov, Dmitry Gaynullin
arXiv Computation and Language
Sep 17

The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models

The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.

By Elzbieta Dawidek (University of Lower Silesia DSW Ideis)
arXiv Computation and Language
Aug 31

Lexically conditioned realization ambiguity in Korean predicate morphology

The article investigates Korean predicate morphology, showing that a sequence of canonical morphemes and grammatical labels does not uniquely determine the surface form for certain predicates. It demonstrates that identical or nearly identical stem-ending configurations can produce different outputs depending on lexical identity and realization class membership. The study frames this as homonymy with inflectional divergence, highlighting that lexical meaning, subcategorization, and semantic role structure are essential for determining the correct surface realization.

By Wonjun Oh, KyungTae Lim, Jungyeul Park
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv Machine Learning
Aug 28

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.

By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv Computation and Language
Sep 7

MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.

By Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha
arXiv Computation and Language
Sep 7

How Do Language Models Represent and Use Phonological Information for Allomorph Selection?

Language models trained on tokenized text still reliably produce morphemes whose form depends on phonology, but it was unclear whether this relies on memorization or rule-like generalization. The study shows that for the English indefinite article a/an, the phonological condition is encoded along a single linear direction in trigger-token embeddings, causally drives article selection in token-level wug tests, and is used by the model to forecast the upcoming trigger token’s phonological feature for article choice. The authors further investigate whether this rule-like generalization extends to allomorph selection in other languages and to explicit phonological judgment, offering a mechanistic account that separates generation-time ability from metalinguistic judgments.

By Sangwoo Kim, Sangah Lee
arXiv Computation and Language
Aug 28

Representing and Parsing Korean Constituency Structure at Different Levels of Granularity

The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.

By Jungyeul Park, KyungTae Lim, Zihao Huang, Eunkyul Leah Jo, Yige Chen, Chulwoo Park