arXiv Computation and Language By Siting Liang, Luca Rippe, Omar Adjali, Daniel Sonntag

Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions

Read the original on arXiv Computation and Language →

The paper investigates how the order of generating explanations—whether a rationale is produced before or after the answer—affects vision‑language reasoning. By conducting controlled experiments on knowledge‑intensive QA, visual entailment, and compositional grounding tasks, the authors show that larger models are required for reliable rationale‑first generation, while answer‑first generation is less susceptible to format errors. The study concludes that explanation ordering, model scale, pre‑training knowledge, fine‑tuning, and task structure jointly influence prediction accuracy and reasoning faithfulness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 17

EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing

EDCT-Bench is a benchmark that evaluates the faithfulness of Vision‑Language Models (VLMs) by using Explanation‑Driven Counterfactual Testing (EDCT). EDCT extracts visual concepts from a model’s natural language explanation, applies minimal verified edits to those concepts, and checks whether the model’s answer and explanation remain consistent with the edited image. The benchmark covers three domains—knowledge‑intensive VQA, safety‑critical driving, and 3D spatial reasoning—and reveals significant faithfulness gaps in current VLMs, while also showing that EDCT‑generated counterfactuals can improve training.

By Sihao Ding, Santosh Vasa, Aditi Ramadwar, Thomas Monninger
arXiv AI
Jun 10

V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

arXiv:2512. 11995v2 Announce Type: replace-cross Abstract: While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle in practice with complex open-ended tasks, which usually require multiple rounds of exploration and reasoning in the visual space.

By Chenrui Fan, Yijun Liang, Shweta Bhardwaj, Kwesi Cobbina, Ming Li, Tianyi Zhou