GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.13225v1 Announce Type: cross Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fail...
arXiv:2608.21832v1 Announce Type: new Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether mo...
arXiv:2608.16081v2 Announce Type: replace Abstract: Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained han...
AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
arXiv:2608. 06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception.