arXiv:2601. 03388v3 Announce Type: replace-cross Abstract: Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain a large number of metaphors.
By Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Zhu
arXiv:2608. 15828v1 Announce Type: cross Abstract: Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge.
By Ana Naveriani, Jakob Suchan, Stefano Zoia, Mehul Bhatt, Antonio Lieto, Gian Luca Pozzato
arXiv:2607. 28683v1 Announce Type: cross Abstract: Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to achieve generalisability across different domains.
By Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Dong, Liming Zhu
MetaHOPE is an error‑severity‑aware annotation framework designed to evaluate how well machine translation (MT) and large language models (LLMs) translate metaphors. The authors applied MetaHOPE to three state‑of‑the‑art systems—GoogleMT, GPT5.4, and Hunyuan‑7b—using two human‑annotated metaphor corpora (VUAMC and PSUCMC) for English‑to‑Chinese and Chinese‑to‑English translation. They also produced a bilingual post‑edited gold reference, creating a new resource for metaphor translation research.
By Jiahui Liang, Lifeng Han
The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.
By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object.
The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.
arXiv:2607. 15847v1 Announce Type: cross Abstract: Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation.
By Suzan Awinat, Alfonso Ortega del Puente
arXiv:2606. 02147v1 Announce Type: cross Abstract: Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.
By Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
PRISM is a modality‑agnostic, category‑theoretic framework that measures and refines multimodal analogies by representing them as explicit relational mappings. It introduces a pullback score to quantify relational alignment and an iterative refinement loop that uses this score as feedback to improve generated images. On the AnaloBench benchmark, PRISM’s pullback score alone achieves 82.5% accuracy, and human evaluations show a 57.65% preference for refined outputs, though refinement may sometimes favor visually crowded compositions.
By Mirella Zeisler, Ojas Shirekar, Mircea Lic\v{a}, Chirag Raman
AraMIP introduces a new guideline for annotating metaphors in Arabic, building upon the established MIPVU framework and tailoring it to Arabic’s linguistic features. The authors distinguish three figurative types—Isti'ara (metaphor), kinaya (metonymy/indirect expression), and tashbih (simile)—and apply the procedure to a pilot dataset of 300 sentences (5,277 words). Their analysis highlights Arabic‑specific challenges such as morphological complexity, inconsistent dictionary sense ordering, and a lack of standardized contextual materials for annotators.
By Mandar Marathe, Manar Ali, Sara Nabhani, Raia Abu Ahmad, Ibrahim Baroud, Omar Momen
What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties.