Hugging Face Trending Papers

Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation

arXiv Computer Vision
Aug 26

Variance-Guided Spatial Attention Fusion for Robust End-to-End Driving under Asymmetric Sensor Degradation

The paper introduces Variance‑Guided Spatial Attention Fusion (VG‑SAF), a method for robust end‑to‑end driving that fuses camera and LiDAR data while handling asymmetric sensor degradation. VG‑SAF uses a physically grounded augmentor to generate dense reliability masks, modality‑specific experts to predict per‑pixel reliability scales, and a hybrid attention mechanism that gates unreliable cells and balances modalities. The approach also includes a Laplace uncertainty head to signal severe or combined sensor failures, and demonstrates improved closed‑loop robustness on the CARLA Longest6 benchmark across various degradation scenarios.

By Weizhi Tao, Zengwang Jin, Xiao Wang, Hailong Huang
arXiv AI
Jun 2

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

arXiv:2606. 01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings.

By Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura
arXiv Computer Vision
4d ago

RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation

RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.

By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu
arXiv AI
Sep 7

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.

By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
arXiv AI
Jul 28

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.

By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.

arXiv Computer Vision
Aug 24

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

The paper introduces a multi‑modal traffic sign detection framework that fuses camera and LiDAR data using an Intensity‑Aware Deformable Fusion module to align retro‑reflective LiDAR cues with visual features. It also presents a dual motion‑model tracker to handle non‑linear perspective changes and a semantic attribute classification pipeline that estimates occlusion, readability, sign embeddedness, and road relevance. Evaluated on a dataset covering more than 60 countries and 2,500 hours of driving, the system achieves an Object Miss Ratio of 0.49% across 221,068 sequences, indicating strong global generalization for autonomous driving.

By Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani