arXiv Computer Vision

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.

arXiv Computer Vision
Sep 4

Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

The paper introduces LaPla, a Vision‑Language‑Action framework that uses a latent‑aligned planning approach to convert discrete semantic reasoning into continuous, physics‑constrained driving actions. It employs a residual VQ‑VAE to encode vehicle kinematics into a structured latent space, then projects multimodal inputs—images, past actions, and text—directly into this latent space, allowing a frozen decoder to generate physically plausible trajectories without quantization errors. Experiments on nuScenes and NVIDIA AlpaSim show LaPla reduces long‑horizon L2 error by 15.52% and improves closed‑loop success rates by 33.34 percentage points while cutting inference latency.

By Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
arXiv Machine Learning
Sep 24

Less Language, More Latents: Annotation-Efficient VLAs for Driving

The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.

By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
arXiv AI
Aug 19

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

LoopVLA introduces a recurrent Vision‑Language‑Action architecture that learns to refine multimodal representations, predict actions, and estimate when further refinement is unnecessary. By iteratively applying a shared Transformer block and producing a sufficiency score at each step, it decouples refinement from fixed layer indices and aligns confidence scores with action quality through a self‑supervised objective. Experiments on LIBERO, LIBERO‑Plus, and VLA‑Arena demonstrate that LoopVLA reduces model parameters by 45% and boosts inference throughput up to 1.7× while matching or surpassing strong baselines in task success.

By Boyang Shen, Kaixiang Yang, Hao Wang, Qiuyu Yu, Qiang Xie, Qiang Li, Zhiwei Wang
Hugging Face Trending Papers
Jul 6

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.

arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
Hugging Face Trending Papers
Aug 13

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling.