LegendBench: A Diagnostic Benchmark for Legend Understanding with Counterfactual Interventions
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...
The paper investigates how vision‑language models (VLMs) read exact values from vertical bar charts. Using controlled counterfactual activation patching on Qwen2.5VL‑7B‑Instruct and InternVL3.5‑8B, the authors find that the bar‑top region contributes more to answer accuracy than the bar body, and that legend and series states lose recoverability earlier than geometry and axis‑scale states. They also show that prompt‑series positions mediate legend influence and that the two models differ in how they combine geometry and scale information from separate image donors.
arXiv:2607. 16311v1 Announce Type: cross Abstract: Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.
arXiv:2607. 22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data.
arXiv:2608. 07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify.
EDCT-Bench is a benchmark that evaluates the faithfulness of Vision‑Language Models (VLMs) by using Explanation‑Driven Counterfactual Testing (EDCT). EDCT extracts visual concepts from a model’s natural language explanation, applies minimal verified edits to those concepts, and checks whether the model’s answer and explanation remain consistent with the edited image. The benchmark covers three domains—knowledge‑intensive VQA, safety‑critical driving, and 3D spatial reasoning—and reveals significant faithfulness gaps in current VLMs, while also showing that EDCT‑generated counterfactuals can improve training.