arXiv AI

Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

arXiv:2606. 24093v1 Announce Type: cross Abstract: We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work.

arXiv Computation and Language
Aug 25

Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

The paper introduces Peony, a benchmark designed to evaluate large language models’ ability to comprehend the ‘poetic logic’ of modern Chinese poetry. It defines this logic through four tasks across stanza, line, and imagery levels and tests six mainstream LLMs under both non‑thinking and thinking configurations. Results show current LLMs struggle with this literary reasoning, highlighting Peony’s role in revealing these limitations.

By Tian Lan, Shanshan Wang, Zehua Duo, Jiang Li, Guanglai Gao, Derek F. Wong, Xiangdong Su
arXiv Computation and Language
Sep 18

Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

Neo-Classic is a new benchmark designed to evaluate linguistic‑aesthetic reasoning in Classical Chinese poetry. It uses an out‑of‑sample dataset of strictly metrical poems written by contemporary experts and a set of reverse‑understanding probes, avoiding reliance on historical corpora. Experiments with leading LLMs show a 20–50% performance drop on contemporary texts and low accuracy (0–13%) on discourse‑level ordering, indicating that current models excel at local patterns but struggle with global hierarchical planning.

By Han Zhang, Zihan Gu, Zhiyuan Wang, Tianyi Ma, Jiacheng Lu, Xinyan Zhang, Yuhao Wei, Cheng Hua
arXiv Computation and Language
Aug 31

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.

By Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka
Hugging Face Trending Papers
Aug 12

JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis.