arXiv Computer Vision

Relationally Grounded Latent World Models for Autonomous Driving

Relationally Grounded Latent World Models for Autonomous Driving proposes using traffic scene graphs as privileged semantic supervision for latent world representations. The approach builds actor‑centric scene graphs from nuScenes 3D annotations, encodes their relational structure with a frozen text embedding model, and aligns visual latent representations to this semantic target during training. At inference the supervision branch is removed, requiring no scene graphs or 3D annotations and adding no extra computation, while achieving a 5.9% reduction in average trajectory L2 error and a 52.4% drop in collision rate compared to the LAW baseline.

arXiv Computer Vision
6d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv Computer Vision
Sep 16

NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving

NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.

By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
arXiv Machine Learning
Sep 24

Less Language, More Latents: Annotation-Efficient VLAs for Driving

The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.

By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
arXiv Computer Vision
Aug 27

Latent Chain-of-Thought World Modeling for End-to-End Driving

Latent-CoT-Drive (LCDrive) is a vision‑language‑action model for autonomous driving that replaces natural‑language chain‑of‑thought reasoning with a latent language capturing possible outcomes of driving actions. The model interleaves action‑proposal tokens, aligned with the model’s output actions, and world‑model tokens grounded in a learned latent world model to reason about future outcomes. After a supervised cold‑start using ground‑truth future rollouts, LCDrive is further refined with closed‑loop reinforcement learning, achieving faster inference, higher‑quality trajectories, and greater gains from interactive RL than both non‑reasoning and text‑reasoning baselines on a large‑scale end‑to‑end driving benchmark.

By Shuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian, Yurong You, Yan Wang, Wenjie Luo, Yulong Cao, Philipp Krahenbuhl, Marco Pavone, Boris Ivanovic
Hugging Face Trending Papers
Jul 6

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.

arXiv AI
Sep 24

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

arXiv:2609. 28366v1 Announce Type: cross Abstract: Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning.

By Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li
arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui