arXiv:2608. 13495v1 Announce Type: cross Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis.
By Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman
MatchFusion is a learnable module that performs explicit-implicit instance matching for spatio‑temporal multimodal autonomous driving. It initializes pairwise affinities with geometric similarity and category consistency, then refines associations using instance embeddings to guide a residual aggregation operator for adaptive information exchange. Experiments on nuScenes show that MatchFusion improves perception accuracy, reduces FLOPs by 55.3% and GPU memory usage by 39.3%, and adds only 3.7% of total perception latency.
By Xiaoyu Li, Jiajia Fu, Long Shi, Tianyu Du, Ruihang Li, Xian Wu, Lijun Zhao, Yingtao Zhang, Lining Sun, Ruifeng Li
arXiv:2606. 20274v1 Announce Type: new Abstract: Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories.
By Shihao Ji, HongXi Li, Zihui Song, Mingyu Li
arXiv:2608.20890v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning,...
By Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang, Yaonan Wang, Ajmal Mian, Mike Zheng Shou
arXiv:2606. 09109v1 Announce Type: cross Abstract: Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking.
By Manyi Yao, Sparsh Garg, Christian Shelton, Amit Roy-Chowdhury, Abhishek Aich
arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.
By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes
arXiv:2607. 23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven.
By Yuhang Wang, Lingyao Li, Hao Zhou
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
DGCPath is a Distribution‑Aware Generative Contrastive framework designed for self‑supervised path representation learning. It combines a diffusion‑based view generator, a variational contrastive mechanism that aligns latent features at the distribution level, and a generative cross‑supervision module for view‑level consistency. Experiments on three real‑world trajectory datasets show that DGCPath surpasses state‑of‑the‑art baselines on two downstream tasks, indicating stronger generalization and representation effectiveness.
By Sean Bin Yang, Hao Miao, Zongyi Xu, Jilin Hu, Xiangmeng Wang, Hua Lu, Bin Yang, Christian S. Jensen
EventVL introduces the first generative event-based multimodal large language model (MLLM) designed for explicit semantic understanding of event streams. The framework leverages a newly annotated dataset of nearly 1.4 million event–image/video–text pairs and incorporates an Event Spatiotemporal Representation to capture comprehensive event information, along with Dynamic Semantic Alignment to refine sparse semantic spaces. Experiments demonstrate that EventVL outperforms existing MLLM baselines in event captioning and scene description generation tasks, advancing the field of event vision.
By Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong
arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.
By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.