arXiv AI By Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao

AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

Read the original on arXiv AI →

AesCanvas is a new dataset and benchmark that evaluates image aesthetic models on two fronts: CritiqueCanvas, which contains 519,136 instruction–response pairs for long‑form, multi‑dimensional critique across photography, painting, and virtual imagery, and ContextCanvas, which offers 301 expert‑reviewed use scenarios to assess contextual aesthetic suitability. The benchmark tests closed‑source, open‑weight general, and aesthetic‑specific multimodal large language models, revealing that models excel at critique generation but lag in context‑sensitive judgment. The study shows that aesthetic specialization does not reliably transfer to contextual suitability and highlights the need for culturally situated, evidence‑grounded suitability as a distinct objective for aesthetic modeling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

The study evaluates whether multimodal large language models (MLLMs) can produce open‑ended aesthetic critiques comparable to humans. Eight open‑weight MLLMs (7 B–397 B) and GPT‑5.5 were tested on 1,227 r/photocritique posts under various prompts, revealing that reference‑based similarity metrics often misrepresent model performance, while shorter critiques and image‑omission had limited impact. Human judges and annotators found the models’ critiques largely different from human ones, noting that models tend to be overly comprehensive and repetitive rather than selective and specific.

By Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv AI
Jul 7

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.

By Yinsheng Yao, Yan Liu, Chen Ye
arXiv AI
Sep 3

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for evaluating visual metaphor generation in text-to-image models, comprising 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. It proposes a hybrid evaluation framework using a multiple-choice question protocol and dimension-based scoring to assess metaphorical fidelity. Experiments on 11 T2I models show that even top proprietary models struggle with compositional structuring and cross-domain mapping, underscoring the need for further research in this area.

By Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
Hugging Face Trending Papers
Sep 2

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

The paper introduces VMetaphor-Bench, a benchmark for assessing visual metaphor generation in text-to-image models. It contains 1,500 curated metaphors across three levels and ten categories, each paired with two prompts of varying specificity. The authors evaluate 11 T2I models using a hybrid MLLM-as-judge framework that combines a large multiple-choice question set with dimension-based scoring, finding that even top proprietary models struggle with compositional structuring and cross-domain mapping.

arXiv AI
Jun 30

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.

By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon