arXiv AI
Sep 1

Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation

Imag‑Eval is a new language‑grounded benchmark for evaluating Text‑to‑Image models, focusing on how well they follow compositional natural‑language instructions. It disentangles prompt length from compositional difficulty by independently varying the number of instances and the combination of constraints (rules), providing 1,140 prompts and 8,842 rule combinations. The study shows that for structured skills, the difficulty is mainly driven by the number of grounded rules and their binding to instances rather than prompt length alone.

By Ibrahim Mohamed Serouis, David Jaramillo Duque
arXiv AI
Aug 26

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.

By Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu
arXiv AI
Aug 14

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
arXiv AI
Aug 26

Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing

The paper introduces QC‑T2I‑Bench, a question‑centric framework that transforms open text‑to‑image prompts into atomic questions and arranges them using Davidsonian Scene Graphs. It employs hierarchy‑constrained aggregation to prune downstream questions when prerequisites fail and to weight simple and complex prompts differently, enabling joint success measurement and comparison of repeated entities across prompts. Evaluations on English and Chinese prompts show that joint completion drops from 80.7% for two‑capability components to 37.2% for seven‑plus components, and the same records are reused for a cost‑aware routing system that achieves ERNIE’s performance with 21.3% fewer GPU‑seconds per million prompts.

By Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian