arXiv AI By Yuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin, Bin Lin, Peng Jin, Jiaqi Liao, Chaoran Feng, Fanqing Meng, Kunpeng Ning, Bin Zhu, Li Yuan

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Read the original on arXiv AI →

arXiv:2503. 07265v4 Announce Type: replace-cross Abstract: Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.

By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick
arXiv Computer Vision
Sep 3

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

The paper introduces TIC‑Bench, a new benchmark for evaluating multimodal large language models on deeply interleaved text‑image contexts. It covers logical, temporal, and spatial association tasks, totaling 2,280 questions across eight specific types. The authors benchmarked ten state‑of‑the‑art MLLMs, finding a significant performance gap versus human experts and highlighting persistent challenges in integrating evidence across interleaved visual and textual inputs.

By Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo
arXiv Computation and Language
Sep 18

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.

By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
arXiv AI
Sep 17

CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art

CompArt introduces a new approach to aesthetic alignment in text-to-image generation by using the Principles of Art (PoA) such as Balance, Rhythm, and Emphasis to define explicit compositional constraints. The authors create a large dataset of 80,032 WikiArt images, each annotated with PoA analyses generated by a multimodal LLM, and present ArtDapter, a lightweight adapter that steers a pretrained diffusion model along ten PoA dimensions while preserving semantic fidelity. Experiments demonstrate that CompArt outperforms strong baselines in adhering to PoA controls under a dual evaluation protocol.

By Zhe Jin, Tat-Seng Chua