arXiv Computation and Language

Lexically conditioned realization ambiguity in Korean predicate morphology

The article investigates Korean predicate morphology, showing that a sequence of canonical morphemes and grammatical labels does not uniquely determine the surface form for certain predicates. It demonstrates that identical or nearly identical stem-ending configurations can produce different outputs depending on lexical identity and realization class membership. The study frames this as homonymy with inflectional divergence, highlighting that lexical meaning, subcategorization, and semantic role structure are essential for determining the correct surface realization.

arXiv Computation and Language
Aug 28

Representing and Parsing Korean Constituency Structure at Different Levels of Granularity

The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.

By Jungyeul Park, KyungTae Lim, Zihao Huang, Eunkyul Leah Jo, Yige Chen, Chulwoo Park
arXiv Computation and Language
4d ago

Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection

The paper investigates Japanese past‑tense verb inflection in neural morphological generation systems, focusing on a rare irregular subtype where stems end in /e/ and require gemination before the suffix. Despite overall high accuracy (>97%), this <1% subclass accounts for 30–43% of residual errors, amplifying its impact by 34–48 times its frequency. Ablation experiments show that removing this specific subtype yields larger accuracy gains than removing all irregular verbs, highlighting the role of low‑frequency patterns and orthographic processes in error concentration.

By Wen Zhang
arXiv AI
Sep 18

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.

By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam
arXiv Computation and Language
Sep 17

The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models

The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.

By Elzbieta Dawidek (University of Lower Silesia DSW Ideis)
arXiv Computation and Language
Aug 27

The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers

The paper studies how transformer representations evolve across layers by examining the intrinsic dimensionality (ID) of token embeddings and their neighborhood structures. It finds that closed‑class tokens expand and collapse earlier than open‑class tokens, and that these changes are linked to shifts in local geometry. The authors compare encoder and decoder models, showing distinct layer‑wise behaviors, and demonstrate that geometric features alone can predict a token’s part‑of‑speech and reveal how semantic content changes across layers.

By Samuele Vallisa, Federico Ravenda, Claudio Palominos, Rui He, Andrea Raballo, Antonietta Mira, Philipp Homan, Wolfram Hinzen
arXiv Computation and Language
Sep 4

Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords

The paper introduces the Fixed Suffix Dependency Ratio (FSDR) as a metric to measure how much loanwords depend on fixed derivational suffixes for gender assignment. Analyzing 1,832 Latvian noun lemmas, it finds that feminine loanwords rely more on fixed suffixes while masculine loanwords are largely free‑choice, a pattern that has intensified in recent usage. FSDR offers a quantitative tool for distinguishing morphological anchoring from default gender in language contact scenarios.

By Yelingyun Zhang, Atis Kapenieks, Marina Platonova