arXiv AI By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

Read the original on arXiv AI →

KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 31

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.

By Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka
arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv Computation and Language
Sep 11

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

CHRONOBERG is a temporally structured corpus of English book texts covering 250 years, curated from Project Gutenberg and enriched with temporal annotations. It enables quantification of lexical semantic change via time‑sensitive Valence‑Arousal‑Dominance analysis and the creation of historically calibrated affective lexicons. Experiments show that language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts, highlighting the need for temporally aware training and evaluation pipelines.

By Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski