arXiv AI

Iterative Visual Thinking and the Self-Correction Mirage in VLM Grounding

arXiv:2606. 13156v2 Announce Type: replace-cross Abstract: Letting a vision-language model (VLM) think longer at test time has driven much recent progress.

arXiv AI
Sep 17

The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

The paper investigates how Vision‑Language Models (VLMs) often report high confidence even after self‑correcting or arriving at wrong answers, a phenomenon the authors attribute to the verbalized confidence being largely independent of the model’s reasoning trajectory. By analyzing content variation, token masking, and hesitation markers, the authors demonstrate that confidence does not adequately reflect the actual reasoning process and that calibration training can sometimes worsen this disconnect. To address this blind spot, they introduce the Trajectory‑Grounding Score (TGS) in two forms—TGS‑self and TGS‑pair—and propose TGS‑Bench, a suite of 10 benchmarks that reveal divergences between conventional calibration metrics and trajectory‑grounded confidence.

By Jisoo Yang, Jaeho Han, Trung X. Pham, Junyeong Kim
arXiv Computation and Language
Sep 1

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

The paper introduces a framework that combines world models, which generate concrete visual rollouts of possible futures, with multimodal large language models (MLLMs) that perform abstract reasoning. It proposes a controlled concrete reasoning approach and a new training method called Privileged‑Future On‑Policy Self‑Distillation (PF‑OPSD), which uses ground‑truth future videos as privileged teacher context during training while the student model never sees true futures at test time. Experiments on two human‑verified benchmarks, VRQABench and OpenWorldQA, show that PF‑OPSD improves performance by about 10–11% over baselines and enhances robustness to noisy or conflicting rollouts.

By Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen
arXiv Machine Learning
1d ago

Same Reward, Different Skills: When Multimodal RL Learns to Look

The paper demonstrates that reinforcement learning with verifiable rewards (RLVR) can improve vision‑language benchmark performance even when models are trained without visual input. When real images are introduced at test time, models trained blind recover about half of the performance gain at 3B parameters and nearly four‑fifths at 7B, but extended real‑image training can erode grounding while benchmark gains persist. The authors propose a visual resolvability rule and show that requiring visual evidence for correct answers leads to significant improvements in target discovery and generalization to unseen question types, while controls confirm that the gains stem from actual visual grounding rather than artifacts.

By Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Yiqiang Chen
arXiv Computer Vision
Sep 16

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.

By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
arXiv Computer Vision
Sep 22

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.

By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
arXiv AI
Aug 7

Visual Grounding in Zero-Shot Vision-Language Control

arXiv:2608. 06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception.

By J. de Curt\`o, Dayani Plasencia, Diego S\'anchez, I. de Zarz\`a