arXiv AI By Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong, Evan Frick, I-Hung Hsu, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Read the original on arXiv AI →

arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Multimodal Language Models as Text-to-Image Model Evaluators

Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.

By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
arXiv Machine Learning
Jun 9

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

arXiv:2601. 04498v2 Announce Type: replace Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information.

By Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen
arXiv Computer Vision
Aug 27

Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric

The paper introduces T23D-CompBench, a new benchmark for fine‑grained text‑to‑3D quality assessment that includes 3,600 textured meshes generated from ten state‑of‑the‑art models and 129,600 human ratings. It also proposes Rank2Score, a two‑stage rank‑learning metric that first trains with supervised contrastive regression and curriculum learning, then refines predictions using mean opinion scores to better align with human judgments. Experiments show Rank2Score outperforms existing metrics and can be used as a reward function for generative model optimization.

By Bingyang Cui, Yujie Zhang, Qi Yang, Zhu Li, Yiling Xu