arXiv AI

Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving

arXiv:2608. 03330v1 Announce Type: new Abstract: This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations.

arXiv Machine Learning
Sep 25

Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity

The paper introduces active client selection strategies for federated learning in autonomous vehicle trajectory prediction, addressing challenges of high scene uncertainty and heterogeneous complexity across different driving environments. It proposes uncertainty-aware selectors that use per-client negative log-likelihood and aleatoric uncertainty, as well as a joint selector that balances scene complexity and uncertainty to prioritize informative clients. Experiments on the Argoverse dataset show that federated models outperform local training, with uncertainty-aware selection speeding convergence and improving key metrics, while the joint selector yields the best generalization under strong heterogeneity.

By Yiming Xie, Muzi Peng, Fei Miao, Ningfang Mi, Lili Su
Hugging Face Trending Papers
Jun 25

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existing predictors excel at minimizing standard displacement errors, they often overlook the adherence to lane topology of multimodal predictions, particularly for lower-probability modes.

arXiv Computer Vision
Aug 25

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM introduces a visual geometry world action model that predicts future scene geometry instead of future images, using point clouds to capture spatial structure and transformations. The model is pretrained to forecast geometry, then a geometry-conditioned action head predicts ego trajectories. Experiments show that this geometry-based approach yields stronger driving policies than image-based alternatives.

By Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman
arXiv Computer Vision
6d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
Hugging Face Trending Papers
Aug 18

Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.

arXiv Computer Vision
Sep 4

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Drive‑HWM introduces a hierarchical slow‑fast world modeling framework for autonomous driving. The slow model predicts multi‑step future representations, while the fast model jointly predicts the next frame and immediate action using a lightweight multimodal backbone and an autoregressive expert. Dynamic‑Aware Latents, learned through optical‑flow prediction, explicitly capture motion dynamics, and experiments on NAVSIM v1 and v2 show strong driving performance with validated ablation studies.

By Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan