arXiv AI By Zhenyu Hu, Qing Wang, Te Cao, Luo Liao, Longfei Lu, Liqun Liu, Shuang Li, Hang Chen, Mengge Xue, Yuan Chen, Chao Deng, Peng Shu, Huan Yu, Jie Jiang

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

Read the original on arXiv AI →

arXiv:2603. 08090v3 Announce Type: replace-cross Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Multimodal Language Models as Text-to-Image Model Evaluators

Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.

By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
arXiv Computer Vision
Sep 4

Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

The paper introduces a benchmark and evaluation system for measuring how well generative image models preserve the identity of a subject across generation, editing, restoration, and multi‑subject scenarios. It compares three paradigms—input context, trainable subject‑specific parameters, and a persistent identity layer—showing that persistent identity consistently improves fidelity while keeping image quality and instruction adherence high. The study finds that identity preservation remains a distinct limitation of current foundation models, especially under iterative edits, small scales, severe degradation, and multi‑subject composition.

By Mengwei Ren, Xuaner Zhang, Zhihao Xia