arXiv:2608. 08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches.
By Pedram MohajerAnsari, Amir Salarpour, Mert D. Pes\'e
arXiv:2603. 09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars.
By Kanishkha Jaisankar, Pranav M. Pawar, Diana Susan Joseph, Raja Muthalagu, Mithun Mukherjee, Dnyaneshawar Mantri, Ramjee Prasad
SMDDFNet is a deep learning detector designed for traffic sign images, addressing challenges such as small objects, scale variation, and occlusion. It combines a Dynamic Dual Fusion (DDF) module—integrating multi-scale attention and frequency‑domain dynamic filtering—with a state‑space modeling backbone that captures long‑range dependencies efficiently. A multi‑scale feature fusion neck further aggregates pyramid features, enabling robust localization of small signs while maintaining real‑time throughput on datasets like TT100K, GTSDB, PASCAL VOC, and Roboflow.
By TianYi Yu, DaJian Zhong, Lilin Wang
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios.
The paper surveys one‑stage object detectors for autonomous driving, covering the evolution from early models like YOLOv1 and SSD to recent real‑time architectures such as YOLOv10 and anchor‑free detectors like FCOS and CenterNet. It compares these methods on design choices, feature‑fusion strategies, loss functions, deployment trade‑offs, and benchmark performance, while also summarizing datasets, evaluation metrics, open challenges, and future research directions. The survey emphasizes how one‑stage detectors balance speed, accuracy, efficiency, and robustness, noting the gap between benchmark results and dependable real‑world performance.
By Jonel Roman, Ryan Sirjue, Peter Nguyen, Daniel Krutky, Juan Jesus, Sudip Dhakal
arXiv:2607. 22714v1 Announce Type: cross Abstract: Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power.
By Sai Sidharth D
The paper introduces a multi‑modal traffic sign detection framework that fuses camera and LiDAR data using an Intensity‑Aware Deformable Fusion module to align retro‑reflective LiDAR cues with visual features. It also presents a dual motion‑model tracker to handle non‑linear perspective changes and a semantic attribute classification pipeline that estimates occlusion, readability, sign embeddedness, and road relevance. Evaluated on a dataset covering more than 60 countries and 2,500 hours of driving, the system achieves an Object Miss Ratio of 0.49% across 221,068 sequences, indicating strong global generalization for autonomous driving.
By Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani
arXiv:2608.31107v1 Announce Type: new
Abstract: The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalizatio...
By Lucas Wojcik, Gabriel E. Lima, Sergio M. Silva Jr., Eduil Nascimento Jr., David Menotti
The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalization power that leverages the immense pre-training k...
arXiv:2606. 05149v1 Announce Type: cross Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injury-risk-relevant categories from naturalistic roadway video do not exist in the open literature.
By Gandhimathi Padmanaban, Fred Feng
The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.
By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv:2609.17134v1 Announce Type: new
Abstract: Neuromorphic vision systems operate under strict constraints on bandwidth, memory, and energy, particularly at the edge, motivating early mechanisms fo...
By Luca Peres, Giulia D'Angelo, Chiara Bartolozzi, Oliver Rhodes