arXiv AI By Jasmine Owers, Edwin Simpson, Martha Lewis

As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language

Read the original on arXiv AI →

arXiv:2606. 18922v1 Announce Type: cross Abstract: Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and spoken language.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

arXiv:2606. 02147v1 Announce Type: cross Abstract: Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.

By Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
arXiv Computation and Language
Sep 4

To What Extent Do Large Language Models Understand Bangla Idioms?

The paper introduces the first large‑scale benchmark dataset of Bangla idioms, along with a synthetic multiple‑choice question set for idiom meaning identification. It evaluates recent large language models on three idiom‑related tasks—paraphrasing, idiom span detection, and meaning identification—using zero‑shot and few‑shot prompting. Results show significant variability across models, with Phi‑4‑mini‑instruct best at paraphrasing, Kimi‑K2‑32b‑instruct excelling at span detection, and Gemini‑2.5‑flash leading in meaning identification.

By Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi, Swakkhar Shatabda
arXiv AI
Sep 15

Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

The paper introduces NegCue, a large-scale dataset of 1.8 million samples that includes single-word, multi-word, and affixal negation cues, totaling over 200 unique forms. The authors pre-train encoder-only language models and large language models on this dataset to study how different negation types influence understanding. Experiments on five downstream benchmarks reveal that affixal negations provide the most significant performance gains, whereas single-word negations yield modest improvements, and that additional pre-training benefits both model types.

By Tian Tan, Eduardo Blanco
arXiv Machine Learning
Aug 5

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.

By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv Computation and Language
Sep 4

Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

The paper explores how large language models (LLMs) can predict typological features using an in-context learning approach with data from URIEL+ and Glottolog. Zero‑shot prompting alone is inadequate, but providing phylogenetic and geographic neighbour evidence enables LLMs to outperform all baselines, even for low‑resource languages. Additionally, most LLM rationales align with the supplied evidence, suggesting a move toward explainable predictions.

By Qianwen Wang, York Hay Ng, Aditya Khan, En-Shiun Annie Lee