Hugging Face Trending Papers

ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation

arXiv Computation and Language
Sep 2

IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals

IntroConformal introduces a training‑free Conformal Risk Control framework that offers finite‑sample, distribution‑free factuality guarantees for Large Vision‑Language Models. It uses introspective signals—layer‑wise semantic stability and verification probability derived from the model’s own hidden states—to assess claim factuality. Experiments across multiple LVLM architectures show that IntroConformal meets the conformal risk guarantee while reducing abstention and matching or surpassing external verifier baselines in claim‑level discrimination.

By Md. Atabuzzaman, Christian Alexander, Chris Thomas
arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv Computer Vision
Sep 18

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.

By Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
arXiv AI
Sep 17

The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

The paper investigates how Vision‑Language Models (VLMs) often report high confidence even after self‑correcting or arriving at wrong answers, a phenomenon the authors attribute to the verbalized confidence being largely independent of the model’s reasoning trajectory. By analyzing content variation, token masking, and hesitation markers, the authors demonstrate that confidence does not adequately reflect the actual reasoning process and that calibration training can sometimes worsen this disconnect. To address this blind spot, they introduce the Trajectory‑Grounding Score (TGS) in two forms—TGS‑self and TGS‑pair—and propose TGS‑Bench, a suite of 10 benchmarks that reveal divergences between conventional calibration metrics and trajectory‑grounded confidence.

By Jisoo Yang, Jaeho Han, Trung X. Pham, Junyeong Kim
arXiv AI
3d ago

AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation

AVERT-VLN introduces a closed‑loop framework for vision‑and‑language navigation that incorporates an abstention‑aware Monitor to detect instruction‑execution inconsistencies. The Monitor is trained on a new LOSTNAV dataset of 20K counterfactual risk trajectories and fine‑tuned on 40K normal trajectories to recognize semantic deviations. During deployment, the Monitor can suspend autonomous navigation and request human guidance, while offline preference learning uses deviation‑associated failures to improve the policy, achieving 76.2% and 66.3% success on R2R‑CE and RxR‑CE unseen splits.

By Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu
arXiv AI
4d ago

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.

By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang
arXiv AI
2d ago

Towards Reliable Vision-Language Models for Autonomous Driving

The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.

By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner