arXiv Computer Vision
Sep 4

FlexMap: Robust HD Map Construction under Flexible Camera Configurations

FlexMap is a vectorized high‑definition map construction framework that works with flexible camera configurations without needing calibrated rigs or explicit 2D‑to‑BEV transformations. It replaces geometric projection with a geometry foundation model that encodes cross‑view 3D structure, and uses a spatial‑temporal enhancement module and a camera‑aware decoder to separate spatial reasoning from temporal aggregation. Experiments on nuScenes and Argoverse 2 show that FlexMap outperforms pose‑dependent baselines and remains accurate even when camera views are missing or pose estimates are inaccurate.

By Run Wang, Chaoyi Zhou, Amir Salarpour, Xi Liu, Zhi-Qi Cheng, Feng Luo, Mert D. Pes\'e, Siyu Huang
arXiv AI
Sep 10

Closed-Loop Evaluation of Bird's-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies

The paper studies how Bird's‑Eye‑View (BEV) maps predicted by Cross‑View Transformers (CVT) can be used directly as inputs to a Behavior‑Cloning (BC) driving policy in the CARLA simulator. It introduces a six‑channel BEV representation and a Kernel Density Estimation (KDE) weighting scheme to focus learning on underrepresented maneuvers. Closed‑loop tests show that the KDE‑weighted model is the only predicted‑BEV agent to finish an episode without infractions, highlighting that global segmentation scores are poor proxies for driving performance and that prediction quality at critical geometries, especially the route channel, is key to reliable navigation.

By Felipe Carlos dos Santos, Eric Antonelo, Gustavo Claudio Karl Couto
arXiv Machine Learning
Jun 3

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

arXiv:2606. 02956v1 Announce Type: cross Abstract: Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity.

By Richard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert, Nils Rack, Kaiwen Wang, Fabian Konstantinidis, Julian Truetsch, Carlos Fernandez, Annika B\"atz, Kevin R\"osch, Marlon Steiner, Willi Poh, Yinzhe Shen, Royden Wagner, Felix Hauser, Dominik Strutz, Jaime Villa, Gleb Stepanov, Holger Caesar, \"Omer \c{S}ahin Ta\c{s}, Frank Bieder, Jan-Hendrik Pauls, Christoph Stiller
arXiv AI
Jun 16

MapDream: Task-Driven Map Learning for Vision-Language Navigation

arXiv:2602. 00222v3 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception.

By Guoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang, Maiyue Chen, Kaihui Wang, Bo Zhang, Zhizhong Su, Deying Li, Zhaoxin Fan