arXiv Machine Learning

Driving on Memory

arXiv Computer Vision
Sep 24

RoadOcc Learns When to Persist, Transport, or Refresh Memory for Roadside Occupancy Prediction

RoadOcc is a new method for roadside occupancy prediction that learns to route information among three memory sources: Persist (fixed-coordinate history), Transport (velocity-addressed history), and Refresh (current evidence). It employs dynamic-aware cross‑attention, multi‑scale voxel velocity estimation, and velocity‑guided dynamic sparse fusion to combine these sources efficiently. On the InfraOcc dataset, RoadOcc achieves 65.29 mIoU and 32.37 dynamic mIoU, outperforming the previous STCOcc baseline by significant margins.

By Xiaokai Bai, Lei Yang, Songkai Wang, Lianqing Zheng, Si-Yuan Cao, Hui-liang Shen
arXiv Computer Vision
22h ago

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

arXiv:2609.38641v1 Announce Type: new Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...

By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
arXiv AI
6d ago

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

By Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll
Hugging Face Trending Papers
Sep 24

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are then refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

arXiv Computer Vision
Aug 31

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.

By Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang