The paper introduces the first large‑scale benchmark dataset of Bangla idioms, along with a synthetic multiple‑choice question set for idiom meaning identification. It evaluates recent large language models on three idiom‑related tasks—paraphrasing, idiom span detection, and meaning identification—using zero‑shot and few‑shot prompting. Results show significant variability across models, with Phi‑4‑mini‑instruct best at paraphrasing, Kimi‑K2‑32b‑instruct excelling at span detection, and Gemini‑2.5‑flash leading in meaning identification.
By Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi, Swakkhar Shatabda
arXiv:2606. 02584v1 Announce Type: cross Abstract: Idiomatic expressions remain a persistent challenge for natural language processing because their meanings are often non-compositional, context-dependent, and difficult to align across languages.
By Ayman Ali Sharara
arXiv:2606. 18922v1 Announce Type: cross Abstract: Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and spoken language.
By Jasmine Owers, Edwin Simpson, Martha Lewis
The paper introduces a new benchmark for Urdu‑to‑English idiomatic translation, featuring 4,000 manually verified sentence pairs in both native Perso‑Arabic script and Romanized Urdu. It evaluates multiple tasks—translation, paraphrasing, idiom span detection, and back‑translation—using various prompting strategies, and finds that state‑of‑the‑art large language models outperform traditional neural machine translation systems, especially in preserving figurative meaning. The study also highlights challenges posed by the lack of standardized orthography in Romanized Urdu, which affects consistency and idiom span detection.
By Muhammad Farmal Khan, Mousumi Akter
arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.
By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv:2606.18389v2 Announce Type: replace
Abstract: Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated dat...
By Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin, Simon Ostermann
The study investigates how multilingual training affects figurative language identification in proverbs, using 742 proverb concepts translated into seven languages. Five models—including multilingual encoders and instruction‑tuned LLMs—were evaluated with varying levels of multilingual supervision, and a new multidimensional annotation framework was introduced to classify proverbs into metaphorical, moral/advisory, cause‑effect, and culture‑specific forms. Results show that adding multilingual data beyond 50% yields limited gains, but the best supervision level depends on the model and language; combining diverse figurative forms improves overall performance, especially for the least frequent culture‑specific form.
By Rama Alomair, Remas Alsubaie, Walaa Saifalislam, Rima Alsonbul, Mona Alnajjar, Razan Aldossari, Haya Alibrahim, Abeer Aldayel
SWORD is a new benchmark that tests large language models’ ability to reject factually incorrect statements across eight major languages by distorting Wikidata triples. The benchmark reveals that models often perform better on semantically plausible distortions than on random ones, indicating a reliance on distributional familiarity rather than true factual verification. It also shows significant performance drops for East Asian languages, with gaps up to 28 percentage points, highlighting asymmetric multilingual factual reasoning capabilities.
By Sanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary.
arXiv:2510. 04120v2 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.
By Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong
MetaHOPE is an error‑severity‑aware annotation framework designed to evaluate how well machine translation (MT) and large language models (LLMs) translate metaphors. The authors applied MetaHOPE to three state‑of‑the‑art systems—GoogleMT, GPT5.4, and Hunyuan‑7b—using two human‑annotated metaphor corpora (VUAMC and PSUCMC) for English‑to‑Chinese and Chinese‑to‑English translation. They also produced a bilingual post‑edited gold reference, creating a new resource for metaphor translation research.
By Jiahui Liang, Lifeng Han
arXiv:2606. 19727v1 Announce Type: cross Abstract: Language models have become essential tools in shaping modern workflows.
By Punit Kumar Singh, Niladri Ghosh, Advait Joshi{\i}nst, Shailee Choudhary, Michael F\"arber, Haiqin Yang