Neo-Classic is a new benchmark designed to evaluate linguistic‑aesthetic reasoning in Classical Chinese poetry. It uses an out‑of‑sample dataset of strictly metrical poems written by contemporary experts and a set of reverse‑understanding probes, avoiding reliance on historical corpora. Experiments with leading LLMs show a 20–50% performance drop on contemporary texts and low accuracy (0–13%) on discourse‑level ordering, indicating that current models excel at local patterns but struggle with global hierarchical planning.
By Han Zhang, Zihan Gu, Zhiyuan Wang, Tianyi Ma, Jiacheng Lu, Xinyan Zhang, Yuhao Wei, Cheng Hua
arXiv:2606. 12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry.
By Haotao Xie
Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited.
arXiv:2608. 11452v1 Announce Type: cross Abstract: Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem.
By Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou
arXiv:2608.23098v1 Announce Type: new
Abstract: Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry...
By Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li, Xin Cong, Yuzhuo Bai, Kangyang Luo, Maosong Sun
arXiv:2606. 24093v1 Announce Type: cross Abstract: We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work.
By Chi-Sheng Chen, Hung-Yun Liu
arXiv:2608.28986v1 Announce Type: new
Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a K...
By Keunhyeung Park, Seunguk Yu, YoungBin Kim
arXiv:2607.00848v4 Announce Type: replace
Abstract: In this opinion paper, we propose MetaHOPE, an error severity-aware annotation framework for evaluating metaphor translations. Metaphors present ch...
By Jiahui Liang, Lifeng Han
arXiv:2609.23951v1 Announce Type: new
Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior w...
By Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe
arXiv:2609.15511v1 Announce Type: cross
Abstract: This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship pe...
By Livia Oddi, Simone Scardapane, Toru Sugimoto, Donatella Genovese
arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.
By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv:2510. 06039v2 Announce Type: replace-cross Abstract: Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language text with verifiable Knowledge Graph (KG) facts.
By Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan, Jiapu Wang