The article investigates Korean predicate morphology, showing that a sequence of canonical morphemes and grammatical labels does not uniquely determine the surface form for certain predicates. It demonstrates that identical or nearly identical stem-ending configurations can produce different outputs depending on lexical identity and realization class membership. The study frames this as homonymy with inflectional divergence, highlighting that lexical meaning, subcategorization, and semantic role structure are essential for determining the correct surface realization.
By Wonjun Oh, KyungTae Lim, Jungyeul Park
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
By Albin Th\"orn Cleland, Eric Cullhed
arXiv:2608. 10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step.
By I\c{s}{\i}l \"Ozg\"u, Yaoxuan Wu, Guy Van den Broeck, Miryung Kim
KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.
By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam
arXiv:2607. 08063v1 Announce Type: cross Abstract: Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone.
By Ryosuke Yamaki, Daichi Mochihashi, Nobutaka Shimada, Tadahiro Taniguchi
arXiv:2607. 10591v2 Announce Type: replace Abstract: Non-binary bottom-up constituency parsing commonly uses arity-specific reductions such as \(\textsc{Reduce-}X\#k\), which jointly specify the mother label and number of children.
By Jungyeul Park, Eunkyul Leah Jo, Zihao Huang
arXiv:2609.15991v1 Announce Type: new
Abstract: Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entr...
By Connor Makowski, Willem Guter
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
The paper introduces discourse dependency (DDP) as a continuous measure of translation difficulty based on how far back a segment must look to resolve references. DDP is computed from named entity re‑mentions and pronominal coreference, and is validated against gold coreference with high reliability. Applying DDP to recent WMT benchmarks reveals a bias toward low‑DDP segments, and experiments show that as DDP increases, no current context‑injection strategy matches human post‑editing quality.
By Ahrii Kim, Chanjun Park, Seong-heum Kim
arXiv:2606. 15521v1 Announce Type: cross Abstract: Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string.
By Kanishk Jain, Matthew Day, Tankut Can
arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.
By Alexandre Bouayad
arXiv:2606. 15144v1 Announce Type: cross Abstract: Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure that underlies word formation.
By Jann Railey Montalan, David Demitri Africa, Jimson Paulo Layacan, Richell Isaiah Flores, Ivan Yuri De Leon, Lance Calvin Gamboa