FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
Read the original on arXiv Computation and Language →FPCO-Dialog is a new benchmark designed to evaluate how vision‑language models correct and cooperate when faced with repeated false premises in multi‑turn dialogues. The dataset contains 1,080 images and 10,800 question turns, organized by visual complexity, object category, and false‑premise class, and follows a 10‑turn protocol where a correct dialogue prefix is followed by repeated false‑premise expressions. Using a model‑agnostic protocol and the CorrTP@K correction‑rate metric, the benchmark reveals significant differences among 20 commercial and open‑source VLMs in their correction tendencies, turn‑wise dynamics, and responses to different false‑premise types.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.