The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.
By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.
ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.
By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2607. 18237v1 Announce Type: cross Abstract: Human visual similarity judgments are context-dependent.
By Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
PRISM is a modality‑agnostic, category‑theoretic framework that measures and refines multimodal analogies by representing them as explicit relational mappings. It introduces a pullback score to quantify relational alignment and an iterative refinement loop that uses this score as feedback to improve generated images. On the AnaloBench benchmark, PRISM’s pullback score alone achieves 82.5% accuracy, and human evaluations show a 57.65% preference for refined outputs, though refinement may sometimes favor visually crowded compositions.
By Mirella Zeisler, Ojas Shirekar, Mircea Lic\v{a}, Chirag Raman
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs).
arXiv:2510. 04120v2 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.
By Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong
arXiv:2608. 11907v1 Announce Type: cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
The paper introduces Visual Metaphor Transfer (VMT), a task that requires models to extract the abstract ‘creative essence’ from a reference image and apply it to a new target subject. It proposes a multi‑agent framework based on Conceptual Blending Theory, using a Schema Grammar to separate relational invariants from visual entities. The system includes perception, transfer, generation, and diagnostic agents, and experimental results show it outperforms state‑of‑the‑art baselines in metaphor consistency, analogy appropriateness, and visual creativity.
By Yu Xu, Yuxin Zhang, Lin Gao, Oliver Deussen, Tong-Yee Lee, Fan Tang
arXiv:2503. 07265v4 Announce Type: replace-cross Abstract: Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content.
By Yuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin, Bin Lin, Peng Jin, Jiaqi Liao, Chaoran Feng, Fanqing Meng, Kunpeng Ning, Bin Zhu, Li Yuan