arXiv AI By TianYi Yu, DaJian Zhong, Lilin Wang

SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection

Read the original on arXiv AI →

SMDDFNet is a deep learning detector designed for traffic sign images, addressing challenges such as small objects, scale variation, and occlusion. It combines a Dynamic Dual Fusion (DDF) module—integrating multi-scale attention and frequency‑domain dynamic filtering—with a state‑space modeling backbone that captures long‑range dependencies efficiently. A multi‑scale feature fusion neck further aggregates pyramid features, enabling robust localization of small signs while maintaining real‑time throughput on datasets like TT100K, GTSDB, PASCAL VOC, and Roboflow.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

One-Stage Object Detectors in Autonomous Driving

The paper surveys one‑stage object detectors for autonomous driving, covering the evolution from early models like YOLOv1 and SSD to recent real‑time architectures such as YOLOv10 and anchor‑free detectors like FCOS and CenterNet. It compares these methods on design choices, feature‑fusion strategies, loss functions, deployment trade‑offs, and benchmark performance, while also summarizing datasets, evaluation metrics, open challenges, and future research directions. The survey emphasizes how one‑stage detectors balance speed, accuracy, efficiency, and robustness, noting the gap between benchmark results and dependable real‑world performance.

By Jonel Roman, Ryan Sirjue, Peter Nguyen, Daniel Krutky, Juan Jesus, Sudip Dhakal
arXiv Computer Vision
Sep 21

Traffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

The paper presents a traffic sign recognition system that extends YOLOv2 with a branched architecture and geometric feature integration. By adding intermediate prediction layers, the model can terminate inference early for easy cases, reducing computation time, while unsupervised Bayesian segmentation supplies geometric templates to improve classification of visually similar signs. Experiments on a combined GTSDB/GTSRB dataset show that the branched model achieves 0.680 mAP in 0.647 s, and adding geometric verification during inference raises mAP to 0.713.

By Arefeh Rezaei
arXiv Machine Learning
Aug 11

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

arXiv:2608. 08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches.

By Pedram MohajerAnsari, Amir Salarpour, Mert D. Pes\'e
arXiv Computer Vision
Aug 25

Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition

The paper introduces an Attention-Driven Complementarity Resampling framework to enhance cross-modality object detection. It employs a shared channel spatial attention mechanism that exchanges semantic masks between modalities, encouraging the backbone to learn generalized features. Additionally, a learnable channel competition module samples and aggregates features channel‑wise, improving robustness and achieving competitive results on multiple datasets.

By Guandi Wang, Ming Li, Yunsen Xing, Junle Liu