arXiv:2609.17499v1 Announce Type: cross
Abstract: Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions...
By Vicky Feliren, A. Taufiq Asyhari, Muhamad Risqi U. Saputra
IntroConformal introduces a training‑free Conformal Risk Control framework that offers finite‑sample, distribution‑free factuality guarantees for Large Vision‑Language Models. It uses introspective signals—layer‑wise semantic stability and verification probability derived from the model’s own hidden states—to assess claim factuality. Experiments across multiple LVLM architectures show that IntroConformal meets the conformal risk guarantee while reducing abstention and matching or surpassing external verifier baselines in claim‑level discrimination.
By Md. Atabuzzaman, Christian Alexander, Chris Thomas
arXiv:2609.24576v1 Announce Type: cross
Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...
By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv:2609.10333v1 Announce Type: new
Abstract: Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-fr...
By Xuan Cuong Ngo, Ngan Le
The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.
By Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
The paper investigates how Vision‑Language Models (VLMs) often report high confidence even after self‑correcting or arriving at wrong answers, a phenomenon the authors attribute to the verbalized confidence being largely independent of the model’s reasoning trajectory. By analyzing content variation, token masking, and hesitation markers, the authors demonstrate that confidence does not adequately reflect the actual reasoning process and that calibration training can sometimes worsen this disconnect. To address this blind spot, they introduce the Trajectory‑Grounding Score (TGS) in two forms—TGS‑self and TGS‑pair—and propose TGS‑Bench, a suite of 10 benchmarks that reveal divergences between conventional calibration metrics and trajectory‑grounded confidence.
By Jisoo Yang, Jaeho Han, Trung X. Pham, Junyeong Kim
Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-free coverage guarantees. However, standard conform...
AVERT-VLN introduces a closed‑loop framework for vision‑and‑language navigation that incorporates an abstention‑aware Monitor to detect instruction‑execution inconsistencies. The Monitor is trained on a new LOSTNAV dataset of 20K counterfactual risk trajectories and fine‑tuned on 40K normal trajectories to recognize semantic deviations. During deployment, the Monitor can suspend autonomous navigation and request human guidance, while offline preference learning uses deviation‑associated failures to improve the policy, achieving 76.2% and 66.3% success on R2R‑CE and RxR‑CE unseen splits.
By Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu
SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.
By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang
The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.
By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv:2608. 19807v1 Announce Type: new Abstract: Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall model performance against annotated ground truth.
By Rongyu Yu, Ke Niu, Fengxiang He
arXiv:2510. 05566v2 Announce Type: replace-cross Abstract: Large language models have achieved impressive performance across diverse tasks.
By Zhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao, Michael von Gablenz