arXiv Machine Learning

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

arXiv:2608. 08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches.

arXiv Computer Vision
Sep 21

Traffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

The paper presents a traffic sign recognition system that extends YOLOv2 with a branched architecture and geometric feature integration. By adding intermediate prediction layers, the model can terminate inference early for easy cases, reducing computation time, while unsupervised Bayesian segmentation supplies geometric templates to improve classification of visually similar signs. Experiments on a combined GTSDB/GTSRB dataset show that the branched model achieves 0.680 mAP in 0.647 s, and adding geometric verification during inference raises mAP to 0.713.

By Arefeh Rezaei
arXiv Computer Vision
Sep 7

FSPGD: Rethinking Black-box Attacks on Semantic Segmentation

FSPGD introduces a feature-space black-box attack for semantic segmentation that targets intermediate representations rather than just output logits. The method uses a dual loss: an external loss to disrupt cross-model feature alignment and an internal loss to reduce consistency among same-class instances. Experiments on Pascal VOC 2012 and Cityscapes show that FSPGD outperforms existing logit-level and segmentation-specific attacks across CNN and Transformer backbones, and its adversarial examples improve robustness when used for training.

By Eun-Sol Park, MiSo Park, Yong-Goo Shin
arXiv AI
Jul 28

Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning

arXiv:2603. 09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars.

By Kanishkha Jaisankar, Pranav M. Pawar, Diana Susan Joseph, Raja Muthalagu, Mithun Mukherjee, Dnyaneshawar Mantri, Ramjee Prasad
arXiv AI
Aug 20

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.

By Ilan Zini, Boussad Addad, Katarzyna Kapusta
Hugging Face Trending Papers
Aug 19

Breaking the weakest link to evade vision language models

The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.

arXiv AI
Sep 24

SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection

SMDDFNet is a deep learning detector designed for traffic sign images, addressing challenges such as small objects, scale variation, and occlusion. It combines a Dynamic Dual Fusion (DDF) module—integrating multi-scale attention and frequency‑domain dynamic filtering—with a state‑space modeling backbone that captures long‑range dependencies efficiently. A multi‑scale feature fusion neck further aggregates pyramid features, enabling robust localization of small signs while maintaining real‑time throughput on datasets like TT100K, GTSDB, PASCAL VOC, and Roboflow.

By TianYi Yu, DaJian Zhong, Lilin Wang