arXiv:2609.12575v1 Announce Type: new
Abstract: Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. Fr...
By Cody Kommers, Mingrui Ye, Evelyn Gius, Daniela Mihai, Hoyt Long, Zheng Yuan, Drew Hemment
The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.
By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
arXiv:2510.26861v4 Announce Type: replace-cross
Abstract: Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practi...
By Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot, Patawee Prakrankamanant, Witthawin Sripheanpol, Pak Lovichit, Sarana Nutanong, Ekapol Chuangsuwanich
arXiv:2605.05593v2 Announce Type: replace
Abstract: Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode...
By Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
arXiv:2608. 00410v2 Announce Type: replace Abstract: Human language is highly polysemous.
By Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths
arXiv:2510. 04391v5 Announce Type: replace Abstract: Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large language model (LLM) populations remains unknown.
By Saurabh Ranjan, Brian Odegaard
The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.
arXiv:2607. 15847v1 Announce Type: cross Abstract: Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation.
By Suzan Awinat, Alfonso Ortega del Puente
arXiv:2508.09852v2 Announce Type: replace-cross
Abstract: How can models help people communicate unusual perceptual experiences without changing what they mean? A recognizable image is only part of t...
By Baihan Lin
arXiv:2609.37230v1 Announce Type: cross
Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a disti...
By Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim
arXiv:2607. 16214v1 Announce Type: cross Abstract: Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear.
By Anna Bavaresco, Ina Klari\'c, Raquel Fern\'andez, Marie-Francine Moens
The Cultural Moment Benchmark (CMB) evaluates video cultural reasoning in Southeast Asia by testing three distinct abilities: naming a cultural concept, visually recognizing it in a video, and locating its sub‑events in time. It contains 306 expert‑curated concepts from seven countries across five categories, with each concept assessed through three stages that use semantic‑similarity distractors, unlabeled video moments, and free‑form temporal localization. Experiments on six vision‑language models reveal varied failure modes, limited cascading between abilities, and differing impacts of audio and subtitles, while a human study shows even experts struggle with concepts from neighboring countries.
By Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo