arXiv:2606. 02147v1 Announce Type: cross Abstract: Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.
By Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
The paper introduces the first large‑scale benchmark dataset of Bangla idioms, along with a synthetic multiple‑choice question set for idiom meaning identification. It evaluates recent large language models on three idiom‑related tasks—paraphrasing, idiom span detection, and meaning identification—using zero‑shot and few‑shot prompting. Results show significant variability across models, with Phi‑4‑mini‑instruct best at paraphrasing, Kimi‑K2‑32b‑instruct excelling at span detection, and Gemini‑2.5‑flash leading in meaning identification.
By Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi, Swakkhar Shatabda
The paper introduces NegCue, a large-scale dataset of 1.8 million samples that includes single-word, multi-word, and affixal negation cues, totaling over 200 unique forms. The authors pre-train encoder-only language models and large language models on this dataset to study how different negation types influence understanding. Experiments on five downstream benchmarks reveal that affixal negations provide the most significant performance gains, whereas single-word negations yield modest improvements, and that additional pre-training benefits both model types.
By Tian Tan, Eduardo Blanco
arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.
By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
The paper explores how large language models (LLMs) can predict typological features using an in-context learning approach with data from URIEL+ and Glottolog. Zero‑shot prompting alone is inadequate, but providing phylogenetic and geographic neighbour evidence enables LLMs to outperform all baselines, even for low‑resource languages. Additionally, most LLM rationales align with the supplied evidence, suggesting a move toward explainable predictions.
By Qianwen Wang, York Hay Ng, Aditya Khan, En-Shiun Annie Lee
arXiv:2510. 04120v2 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.
By Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong
arXiv:2602. 05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.
By Bingru Li
The paper introduces a scope‑conditioned generation framework that incorporates structured stereotype characteristics into prompts for large language models, aiming to improve the quality of counterspeech against online hate speech. The authors validate the method on a new, human‑curated dataset in English, Italian, and Spanish, showing significant gains over generic baselines in factuality, specificity, cogency, and effectiveness for both explicit and implicit stereotypes.
By Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso, Serena Villata
The study investigates how multilingual training affects figurative language identification in proverbs, using 742 proverb concepts translated into seven languages. Five models—including multilingual encoders and instruction‑tuned LLMs—were evaluated with varying levels of multilingual supervision, and a new multidimensional annotation framework was introduced to classify proverbs into metaphorical, moral/advisory, cause‑effect, and culture‑specific forms. Results show that adding multilingual data beyond 50% yields limited gains, but the best supervision level depends on the model and language; combining diverse figurative forms improves overall performance, especially for the least frequent culture‑specific form.
By Rama Alomair, Remas Alsubaie, Walaa Saifalislam, Rima Alsonbul, Mona Alnajjar, Razan Aldossari, Haya Alibrahim, Abeer Aldayel
The paper explores how large language models (LLMs) can predict typological features using an in-context learning approach with data from URIEL+ and Glottolog. Zero‑shot prompting alone is inadequate, but providing phylogenetic and geographic neighbour evidence enables LLMs to outperform all baselines, even for low‑resource languages. Additionally, most LLM rationales align with the supplied evidence, suggesting a move toward explainable typological predictions.
arXiv:2601. 03388v3 Announce Type: replace-cross Abstract: Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain a large number of metaphors.
By Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Zhu
arXiv:2607. 12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation.
By Chun-Yi Kuan, Hung-yi Lee