Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.06419v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-base...
arXiv:2606. 20244v1 Announce Type: cross Abstract: Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact.
The paper investigates how Vision‑Language Models (VLMs) often report high confidence even after self‑correcting or arriving at wrong answers, a phenomenon the authors attribute to the verbalized confidence being largely independent of the model’s reasoning trajectory. By analyzing content variation, token masking, and hesitation markers, the authors demonstrate that confidence does not adequately reflect the actual reasoning process and that calibration training can sometimes worsen this disconnect. To address this blind spot, they introduce the Trajectory‑Grounding Score (TGS) in two forms—TGS‑self and TGS‑pair—and propose TGS‑Bench, a suite of 10 benchmarks that reveal divergences between conventional calibration metrics and trajectory‑grounded confidence.
arXiv:2609.13288v1 Announce Type: new Abstract: Video-language models can answer multiple-choice questions with high confidence yet be wrong. We study whether answer-level reliability scores can be i...
arXiv:2608. 13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain.
arXiv:2607. 08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution.