FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
FPCO-Dialog is a new benchmark designed to evaluate how vision‑language models correct and cooperate when faced with repeated false premises in multi‑turn dialogues. The dataset contains 1,080 images and 10,800 question turns, organized by visual complexity, object category, and false‑premise class, and follows a 10‑turn protocol where a correct dialogue prefix is followed by repeated false‑premise expressions. Using a model‑agnostic protocol and the CorrTP@K correction‑rate metric, the benchmark reveals significant differences among 20 commercial and open‑source VLMs in their correction tendencies, turn‑wise dynamics, and responses to different false‑premise types.
Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action semantics depend on coordinated roles and spatial-te...
arXiv:2607. 16311v1 Announce Type: cross Abstract: Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.
The paper introduces DoublesEval, a diagnostic framework that uses professional doubles badminton to test visual‑language models’ ability to reason about dynamic multi‑agent interactions. It decomposes rallies into key moments and evaluates models across four dimensions—atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction—highlighting specific reasoning failures. The authors also propose TacticCheck, a lightweight consistency checker that improves performance without retraining the models, yet significant gaps remain in tactical reasoning.
ReViCo (Real Visual Correction) is a new benchmark that tests Vision Language Models (VLMs) on the task of correcting text errors in real‑world images, requiring deep understanding of visual text and its context. The study evaluates VLMs using both prompt‑based and targeted training approaches, revealing a significant performance gap between current models and humans. The results show that most VLMs struggle to accurately perceive visual text, leading to frequent correction mistakes, thereby underscoring the need for more robust, text‑aware VLMs.
arXiv:2606. 06890v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence.