arXiv AI

AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

AesCanvas is a new dataset and benchmark that evaluates image aesthetic models on two fronts: CritiqueCanvas, which contains 519,136 instruction–response pairs for long‑form, multi‑dimensional critique across photography, painting, and virtual imagery, and ContextCanvas, which offers 301 expert‑reviewed use scenarios to assess contextual aesthetic suitability. The benchmark tests closed‑source, open‑weight general, and aesthetic‑specific multimodal large language models, revealing that models excel at critique generation but lag in context‑sensitive judgment. The study shows that aesthetic specialization does not reliably transfer to contextual suitability and highlights the need for culturally situated, evidence‑grounded suitability as a distinct objective for aesthetic modeling.

arXiv Computation and Language
Sep 1

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

The study evaluates whether multimodal large language models (MLLMs) can produce open‑ended aesthetic critiques comparable to humans. Eight open‑weight MLLMs (7 B–397 B) and GPT‑5.5 were tested on 1,227 r/photocritique posts under various prompts, revealing that reference‑based similarity metrics often misrepresent model performance, while shorter critiques and image‑omission had limited impact. Human judges and annotators found the models’ critiques largely different from human ones, noting that models tend to be overly comprehensive and repetitive rather than selective and specific.

By Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv AI
Jul 7

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.

By Yinsheng Yao, Yan Liu, Chen Ye
arXiv AI
Sep 3

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.

By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
Hugging Face Trending Papers
Sep 2

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.

arXiv AI
Jun 30

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.

By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
arXiv AI
Sep 17

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.

By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
Hugging Face Trending Papers
Jul 23

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.

arXiv AI
Jun 3

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

arXiv:2605. 20731v2 Announce Type: replace-cross Abstract: Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison.

By Haonan Zhu, Elad Hirsch, Alexandria Minetti, Allison Nulty, Purvanshi Mehta