Hugging Face Trending Papers

Taming Perception Jitter: Uncertainty-Aware LiDAR Object Detection for Reliable Motion Classification

Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner interventions. Unstable bounding box predictions can lead to spurious velocity estimates in tracking and falsely predicted trajectories.

arXiv AI
Jun 2

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

arXiv:2606. 01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings.

By Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura
arXiv Computer Vision
Aug 31

Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation

The paper introduces OccLinker, a lightweight plugin for vision‑based occupancy networks that reduces flickering by efficiently merging historical static and motion cues with current features via a dual cross‑attention mechanism. It generates correction components to refine base network predictions and proposes a new temporal consistency metric to quantify flickering. Experiments on two benchmark datasets show that OccLinker improves performance with minimal computational overhead while effectively diminishing flickering artifacts.

By Fengcheng Yu, Haoran Xu, Canming Xia, Ziyang Zong, Guang Tan
arXiv AI
Sep 10

A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing

The paper introduces a multi‑modal late‑fusion perception pipeline for object detection and tracking in autonomous racing. It combines independent detections from cameras, LiDARs, and RADARs to produce timely and robust state estimates of surrounding vehicles. The tracking framework compensates for detection delays and incorporates vehicle dynamics and track layout knowledge, and its effectiveness is confirmed through real‑world experiments in diverse critical scenarios.

By Davide Malvezzi, Michele Pestarino, Vittoria Cavicchioli, Valentina La Gamba, Silvia Severi, Fabio Bagni, Luca Bartoli, Massimiliano Bosi, Francesco Gatti, Micaela Verucchi, Ayoub Raji, Marko Bertogna
arXiv Computer Vision
Sep 7

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

The paper introduces FlexDepth, a family of self‑supervised monocular depth estimation models designed for robust driving perception. FlexDepth uses a two‑stage static‑dynamic decoupled training strategy and a Scale‑Driven Decoder that selects components based on scale size, enabling efficient feature fusion and high‑precision depth maps. Experiments on driving benchmarks show state‑of‑the‑art performance across arbitrary scales with minimal computational cost, with the smallest model (Flex‑Nano) achieving 37.6 FPS on mobile devices.

By Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan
arXiv AI
Sep 25

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.

By Yuting Zhao, Ziyi Zheng, Shuxiao Li