NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.
Imag‑Eval is a new language‑grounded benchmark for evaluating Text‑to‑Image models, focusing on how well they follow compositional natural‑language instructions. It disentangles prompt length from compositional difficulty by independently varying the number of instances and the combination of constraints (rules), providing 1,140 prompts and 8,842 rule combinations. The study shows that for structured skills, the difficulty is mainly driven by the number of grounded rules and their binding to instances rather than prompt length alone.
The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.
arXiv:2609.00232v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmar...
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
The paper introduces QC‑T2I‑Bench, a question‑centric framework that transforms open text‑to‑image prompts into atomic questions and arranges them using Davidsonian Scene Graphs. It employs hierarchy‑constrained aggregation to prune downstream questions when prerequisites fail and to weight simple and complex prompts differently, enabling joint success measurement and comparison of repeated entities across prompts. Evaluations on English and Chinese prompts show that joint completion drops from 80.7% for two‑capability components to 37.2% for seven‑plus components, and the same records are reused for a cost‑aware routing system that achieves ERNIE’s performance with 21.3% fewer GPU‑seconds per million prompts.