arXiv AI By Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

Read the original on arXiv AI →

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 24

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are then refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

arXiv Computer Vision
Sep 18

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

MM-Future is a world-action model for autonomous driving that generates multiple paired scene-action hypotheses and captures bidirectional interaction within each pair. It initializes each hypothesis from a structured action prior and an independent future scene source, then co-evolves them using a modality-aware diffusion Transformer. The model compresses multi-view video into planning-oriented MM-Tokens and uses a future-conditioned proposal scorer to rank trajectory candidates, achieving strong performance on NAVSIM and HUGSIM benchmarks.

By Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren
arXiv AI
Sep 21

Visual Navigation Transformer with Pose Attention

arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when...

By Beiming Li, Jaime Romero, Jonathan Diller, Vijay Kumar, Alejandro Ribeiro