arXiv:2609.36628v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, partic...
By Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong, Chenwei Xoe, Zhaohe Liao
DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.
By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
arXiv:2609.24128v1 Announce Type: cross
Abstract: Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effect...
By You Liu, Yue Liu, Quanchao Lu, Nick Shipilov
arXiv:2603.07119v3 Announce Type: replace
Abstract: Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall,...
By Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin
AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.
By Min Wang, Ata Mahjoubfar
arXiv:2609.36598v1 Announce Type: new
Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
By Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang