ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decision...
IntroConformal introduces a training‑free Conformal Risk Control framework that offers finite‑sample, distribution‑free factuality guarantees for Large Vision‑Language Models. It uses introspective signals—layer‑wise semantic stability and verification probability derived from the model’s own hidden states—to assess claim factuality. Experiments across multiple LVLM architectures show that IntroConformal meets the conformal risk guarantee while reducing abstention and matching or surpassing external verifier baselines in claim‑level discrimination.
arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...
arXiv:2609.10333v1 Announce Type: new Abstract: Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-fr...
arXiv:2510. 05566v2 Announce Type: replace-cross Abstract: Large language models have achieved impressive performance across diverse tasks.
The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.