arXiv AI

Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning

arXiv:2603. 09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars.

arXiv Computer Vision
Sep 21

Traffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

The paper presents a traffic sign recognition system that extends YOLOv2 with a branched architecture and geometric feature integration. By adding intermediate prediction layers, the model can terminate inference early for easy cases, reducing computation time, while unsupervised Bayesian segmentation supplies geometric templates to improve classification of visually similar signs. Experiments on a combined GTSDB/GTSRB dataset show that the branched model achieves 0.680 mAP in 0.647 s, and adding geometric verification during inference raises mAP to 0.713.

By Arefeh Rezaei
arXiv Machine Learning
Aug 11

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

arXiv:2608. 08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches.

By Pedram MohajerAnsari, Amir Salarpour, Mert D. Pes\'e
arXiv AI
Jun 10

TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

arXiv:2602. 23499v4 Announce Type: replace-cross Abstract: Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render the entire dataset unusable.

By Tugrul Gorgulu, Atakan Dag, M. Esat Kalfaoglu, Halil Ibrahim Kuru, Baris Can Cam, Halil Ibrahim Ozturk, Ozsel Kilinc
arXiv AI
Sep 15

Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation

The paper presents an end‑to‑end system that converts driving footage into dynamic vision sensor (DVS) event streams, augments training with simulated DVS data, and trains a convolutional spiking neural network (Conv‑SNN) to classify pedestrian crossing intent as crossing or non‑crossing. The Conv‑SNN, trained with a class‑balanced loss and surrogate‑gradient learning, achieves high accuracy and F1 scores on JAAD and CARLA DVS datasets, outperforming or matching prior frame‑based methods while operating on sparse temporal representations. The study details architectural choices, neuron dynamics, and training protocols, and provides a convergence analysis and domain‑transfer evaluation.

By Henok Teklu, Mustafa Sakhai, Maciej Wielgosz, Matej Mertik
arXiv Computer Vision
Aug 24

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

The paper introduces a multi‑modal traffic sign detection framework that fuses camera and LiDAR data using an Intensity‑Aware Deformable Fusion module to align retro‑reflective LiDAR cues with visual features. It also presents a dual motion‑model tracker to handle non‑linear perspective changes and a semantic attribute classification pipeline that estimates occlusion, readability, sign embeddedness, and road relevance. Evaluated on a dataset covering more than 60 countries and 2,500 hours of driving, the system achieves an Object Miss Ratio of 0.49% across 221,068 sequences, indicating strong global generalization for autonomous driving.

By Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani
Hugging Face Trending Papers
Aug 18

Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.

arXiv Computer Vision
Sep 4

An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

The paper proposes an ensemble-based self‑taught learning framework for parking space classification that uses unsupervised convolutional autoencoders to learn transferable visual representations from unlabeled data. These learned encoders serve as fixed feature extractors for supervised classification with limited annotated samples, and an ensemble of heterogeneous autoencoders with independent classifier heads is employed to enhance robustness and reduce architectural bias. Experiments on PKLot and CNRPark benchmarks demonstrate that this approach significantly lowers annotation requirements while achieving high accuracies (93–96%) under cross‑dataset evaluation protocols.

By Lucas de Oliveira Cunha, Joelton Deonei Gotz, Paulo Lisboa de Almeida, Andre Gustavo Hochuli
arXiv AI
Aug 20

One-Stage Object Detectors in Autonomous Driving

The paper surveys one‑stage object detectors for autonomous driving, covering the evolution from early models like YOLOv1 and SSD to recent real‑time architectures such as YOLOv10 and anchor‑free detectors like FCOS and CenterNet. It compares these methods on design choices, feature‑fusion strategies, loss functions, deployment trade‑offs, and benchmark performance, while also summarizing datasets, evaluation metrics, open challenges, and future research directions. The survey emphasizes how one‑stage detectors balance speed, accuracy, efficiency, and robustness, noting the gap between benchmark results and dependable real‑world performance.

By Jonel Roman, Ryan Sirjue, Peter Nguyen, Daniel Krutky, Juan Jesus, Sudip Dhakal
Hugging Face Trending Papers
Jul 29

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios.