MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2603. 06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios.
FlexMap is a vectorized high‑definition map construction framework that works with flexible camera configurations without needing calibrated rigs or explicit 2D‑to‑BEV transformations. It replaces geometric projection with a geometry foundation model that encodes cross‑view 3D structure, and uses a spatial‑temporal enhancement module and a camera‑aware decoder to separate spatial reasoning from temporal aggregation. Experiments on nuScenes and Argoverse 2 show that FlexMap outperforms pose‑dependent baselines and remains accurate even when camera views are missing or pose estimates are inaccurate.
arXiv:2605. 06317v4 Announce Type: replace-cross Abstract: Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency.
arXiv:2602. 00222v3 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception.
The paper presents a framework that builds a static point cloud prior map from past camera traversals, augmenting each point with DINOv3 semantic features. During runtime, a local prior patch is retrieved, encoded with a sparse voxel backbone, and fused with lifted multi‑view camera features in bird’s‑eye view. This fused representation is then used by sparse transformer heads to predict 3D objects and vectorized map elements, achieving improved performance on Argoverse 2 without requiring LiDAR for prior‑map construction or online inference.
arXiv:2607.09086v2 Announce Type: replace Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...