arXiv AI

Perceptual Reality Transformer: What Must an Illustration Preserve?

arXiv AI
Sep 3

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.

By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
Hugging Face Trending Papers
Aug 11

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence.

Hugging Face Trending Papers
Sep 2

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.

arXiv Computer Vision
Aug 27

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.

By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson
arXiv Computer Vision
Sep 3

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.

By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo