Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Read the original on arXiv AI →The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior. whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.