arXiv Computation and Language

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.

arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv AI
Jul 16

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

arXiv:2601. 02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora.

By Amirali Ebrahimzadeh, Seyyed M. Salili
arXiv Computation and Language
Sep 4

To What Extent Do Large Language Models Understand Bangla Idioms?

The paper introduces the first large‑scale benchmark dataset of Bangla idioms, along with a synthetic multiple‑choice question set for idiom meaning identification. It evaluates recent large language models on three idiom‑related tasks—paraphrasing, idiom span detection, and meaning identification—using zero‑shot and few‑shot prompting. Results show significant variability across models, with Phi‑4‑mini‑instruct best at paraphrasing, Kimi‑K2‑32b‑instruct excelling at span detection, and Gemini‑2.5‑flash leading in meaning identification.

By Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi, Swakkhar Shatabda