arXiv AI

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

arXiv:2605. 28215v2 Announce Type: replace Abstract: In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples.

arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv Computation and Language
Sep 25

Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions

The paper investigates how the order of generating explanations—whether a rationale is produced before or after the answer—affects vision‑language reasoning. By conducting controlled experiments on knowledge‑intensive QA, visual entailment, and compositional grounding tasks, the authors show that larger models are required for reliable rationale‑first generation, while answer‑first generation is less susceptible to format errors. The study concludes that explanation ordering, model scale, pre‑training knowledge, fine‑tuning, and task structure jointly influence prediction accuracy and reasoning faithfulness.

By Siting Liang, Luca Rippe, Omar Adjali, Daniel Sonntag