arXiv Computation and Language

To What Extent Do Large Language Models Understand Bangla Idioms?

The paper introduces the first large‑scale benchmark dataset of Bangla idioms, along with a synthetic multiple‑choice question set for idiom meaning identification. It evaluates recent large language models on three idiom‑related tasks—paraphrasing, idiom span detection, and meaning identification—using zero‑shot and few‑shot prompting. Results show significant variability across models, with Phi‑4‑mini‑instruct best at paraphrasing, Kimi‑K2‑32b‑instruct excelling at span detection, and Gemini‑2.5‑flash leading in meaning identification.

arXiv AI
Jun 2

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

arXiv:2606. 02147v1 Announce Type: cross Abstract: Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.

By Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
arXiv Computation and Language
Sep 4

Evaluating Large Language Models on Urdu Idioms

The paper introduces a new benchmark for Urdu‑to‑English idiomatic translation, featuring 4,000 manually verified sentence pairs in both native Perso‑Arabic script and Romanized Urdu. It evaluates multiple tasks—translation, paraphrasing, idiom span detection, and back‑translation—using various prompting strategies, and finds that state‑of‑the‑art large language models outperform traditional neural machine translation systems, especially in preserving figurative meaning. The study also highlights challenges posed by the lack of standardized orthography in Romanized Urdu, which affects consistency and idiom span detection.

By Muhammad Farmal Khan, Mousumi Akter
Hugging Face Trending Papers
Jun 2

From Script to Semantics: Prompting Strategies for African NLI

Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning. We present a systematic study of prompting strategies for Natural Language Inference (NLI) in Swahili, Yoruba, and Hausa using the AfriXNLI benchmark.

arXiv AI
Aug 6

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

arXiv:2608. 04670v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions.

By Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni
arXiv Computation and Language
2d ago

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

SWORD is a new benchmark that tests large language models’ ability to reject factually incorrect statements across eight major languages by distorting Wikidata triples. The benchmark reveals that models often perform better on semantically plausible distortions than on random ones, indicating a reliance on distributional familiarity rather than true factual verification. It also shows significant performance drops for East Asian languages, with gaps up to 28 percentage points, highlighting asymmetric multilingual factual reasoning capabilities.

By Sanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun
arXiv Machine Learning
Aug 5

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.

By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang