Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces a classifier‑free method for generating visual counterfactual explanations (VCEs) using Contrastive Analysis (CA). By separating generative factors common to two datasets from those specific to each class, the approach swaps only the salient factors to produce counterfactual images, thereby avoiding reliance on classifier decision boundaries. Leveraging StyleGAN2’s high‑quality synthesis and a feature‑space latent representation, the method supports multiple salient factors per dataset and achieves superior counterfactual quality on three medical imaging datasets.
arXiv:2607. 22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?
arXiv:2608.20621v1 Announce Type: new Abstract: Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coa...
arXiv:2607. 00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts.
The paper introduces visual adaptations of counterfactual tests—vCT and vCCT—to evaluate whether chain-of-thought explanations in vision‑language models faithfully reflect the visual evidence driving predictions. Using these tests, the authors benchmark eight open‑source VLMs on two datasets and find that CoTs often fail to track visual evidence, sometimes omitting removed objects or mentioning them inconsistently. They also release two new datasets, Counter‑SNLI‑VE and Counter‑A‑OKVQA, consisting of image pairs that differ by a single object to facilitate further research.
arXiv:2606. 10571v1 Announce Type: cross Abstract: Adversarial examples reveal vulnerabilities in Vision-Language Pre-training (VLP) models and provide insights for improving robustness.