arXiv AI

SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving

arXiv:2607. 09741v1 Announce Type: cross Abstract: Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents.

arXiv Machine Learning
Jul 1

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

arXiv:2510. 14819v3 Announce Type: replace-cross Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time estimation, mobility prediction, and trajectory similarity analysis.

By Ji Cao, Yu Wang, Tongya Zheng, Jie Song, Qinghong Guo, Zujie Ren, Canghong Jin, Gang Chen, Mingli Song
arXiv Machine Learning
Sep 25

Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity

The paper introduces active client selection strategies for federated learning in autonomous vehicle trajectory prediction, addressing challenges of high scene uncertainty and heterogeneous complexity across different driving environments. It proposes uncertainty-aware selectors that use per-client negative log-likelihood and aleatoric uncertainty, as well as a joint selector that balances scene complexity and uncertainty to prioritize informative clients. Experiments on the Argoverse dataset show that federated models outperform local training, with uncertainty-aware selection speeding convergence and improving key metrics, while the joint selector yields the best generalization under strong heterogeneity.

By Yiming Xie, Muzi Peng, Fei Miao, Ningfang Mi, Lili Su
arXiv Computer Vision
Sep 4

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Drive‑HWM introduces a hierarchical slow‑fast world modeling framework for autonomous driving. The slow model predicts multi‑step future representations, while the fast model jointly predicts the next frame and immediate action using a lightweight multimodal backbone and an autoregressive expert. Dynamic‑Aware Latents, learned through optical‑flow prediction, explicitly capture motion dynamics, and experiments on NAVSIM v1 and v2 show strong driving performance with validated ablation studies.

By Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan
arXiv Computer Vision
6d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv Machine Learning
Aug 19

General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting

The paper introduces a spatio‑temporal traffic forecasting framework that fuses Graph Neural Networks with semantic knowledge from general-purpose knowledge graphs such as Wikidata. By generating embeddings that capture relationships like nearby points of interest, administrative hierarchies, and functional roles of locations, the framework creates additional adjacency matrices that enrich the sensor graph beyond physical connectivity. Experiments with established forecasting methods demonstrate that this external knowledge improves prediction accuracy and offers a path toward better interpretability.

By Mattis thor Straten, Yannick Wolker, Steffen Strohm, Prathvish Mithare, Ralf Krestel, Matthias Renz
Hugging Face Trending Papers
Aug 13

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling.