arXiv Computation and Language

Wiktionary as a Crowdsourced Lexicon for English Dialects

arXiv Computation and Language
2d ago

Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries

The paper tackles the problem of automatically generating dictionary definitions for learner’s dictionaries, focusing on simplicity and clarity. It introduces a new evaluation framework that uses large language models as judges, validated against human annotators with comparable agreement levels. The authors also present an iterative simplification approach that produces definitions scoring highly on their criteria and exhibiting lexical simplicity.

By Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay, Taro Watanabe
arXiv Computation and Language
Sep 22

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.

By Simon Malan, Danel Slabbert, Herman Kamper
arXiv AI
Sep 18

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.

By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam