arXiv:2609.21000v1 Announce Type: cross
Abstract: Spinning frequency-modulated continuous-wave (FMCW) radars have been gaining popularity in autonomous vehicle perception on account of their robustne...
By Eric Xie, Daniil Lisus, Timothy D. Barfoot
arXiv:2606. 05149v1 Announce Type: cross Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injury-risk-relevant categories from naturalistic roadway video do not exist in the open literature.
By Gandhimathi Padmanaban, Fred Feng
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios.
arXiv:2609.22762v1 Announce Type: new
Abstract: Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expe...
By Fengcheng Yu, Dhruv Parikh, Junjie Ye, Maulik Bhatt, Thang Vu, Igor Vasiljevic, Vitor Guizilini, Yue Wang
Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.
By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv:2608.30657v1 Announce Type: new
Abstract: Fixed-viewpoint infrastructure sensors repeatedly observe the same traffic space, making roadside 3D occupancy structurally different from ego-vehicle...
By Lei Yang, Xiaokai Bai, Boqi Li, Chunmian Lin, Li Wang, Ziying Song, Jiahuan Zhang, Enhui Ma, Haibao Yu, Jiaqi Ma, Kaicheng Yu
arXiv:2608. 11770v1 Announce Type: cross Abstract: Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis.
By Vaishnav Raju
RoadOcc is a new method for roadside occupancy prediction that learns to route information among three memory sources: Persist (fixed-coordinate history), Transport (velocity-addressed history), and Refresh (current evidence). It employs dynamic-aware cross‑attention, multi‑scale voxel velocity estimation, and velocity‑guided dynamic sparse fusion to combine these sources efficiently. On the InfraOcc dataset, RoadOcc achieves 65.29 mIoU and 32.37 dynamic mIoU, outperforming the previous STCOcc baseline by significant margins.
By Xiaokai Bai, Lei Yang, Songkai Wang, Lianqing Zheng, Si-Yuan Cao, Hui-liang Shen
arXiv:2608.22086v1 Announce Type: cross
Abstract: We present SweepLSD, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the...
By Yoshiyasu Shimizu
arXiv:2608. 07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels.
By Felix Schaller
We present SweepLSD, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the scan line. Every stage, including connected-compon...
FLINT is a lightweight traversability estimator that uses a 21.6‑million‑parameter backbone—38 times smaller than comparable foundation models—to predict traversability from a single RGB camera. It achieves higher accuracy on held‑out terrain probes and runs at 14.7 FPS on CPU, outperforming a deployed foundation‑model system (WildOS) on 23 of 24 field logs. In closed‑loop field trials, FLINT’s best self‑supervised head reached 99% autonomy, surpassing a human‑label‑trained baseline on the same course.
By William Bonilla, Maxime Boisvert, David-Alexandre Poissant, David Meger, Louis Petit