arXiv Computation and Language By Chaoyi Wu, Yu-Hsiang Tseng, R. Harald Baayen

Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics

Read the original on arXiv Computation and Language →

The paper investigates Mandarin Chinese reduplicative constructions that repeat two-character base words or their constituents, such as expressions meaning ‘in good health’ or ‘discuss a bit’. Using Tencent word embeddings, the study demonstrates that distributional semantics can recover known semantic and grammatical properties of these reduplications, revealing clear semantic and pragmatic differentiation between the two patterns. Procrustes analysis shows that the overall organization of the base-word space is largely preserved in the reduplication space, with local mismatches indicating discourse-pragmatic reorganization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 31

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.

By Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka
arXiv Machine Learning
Aug 5

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.

By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang