Hugging Face Trending Papers

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

Read the original on Hugging Face Trending Papers →

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are then refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
6d ago

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

By Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll
arXiv Computer Vision
Sep 18

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

MM-Future is a world-action model for autonomous driving that generates multiple paired scene-action hypotheses and captures bidirectional interaction within each pair. It initializes each hypothesis from a structured action prior and an independent future scene source, then co-evolves them using a modality-aware diffusion Transformer. The model compresses multi-view video into planning-oriented MM-Tokens and uses a future-conditioned proposal scorer to rank trajectory candidates, achieving strong performance on NAVSIM and HUGSIM benchmarks.

By Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren