Hugging Face Trending Papers

DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild

Read the original on Hugging Face Trending Papers →

Existing object-aware SLAM systems force a trade-off between real-time performance, multi-class support, and the generation of high-fidelity, semantically coherent object models. To address this trade-off, we present DSP-SLAM++, which extends the DSP-SLAM framework with an asynchronous mapping pipeline for real-time performance and dedicated sensor fusion adaptations for a monocular fisheye-LiDAR suite.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
Jul 8

Dynamic Object Detection and Tracking in Construction: A Fisheye Camera and LiDAR Sensor Fusion Model

Robust dynamic object detection and tracking are essential for enabling robots to operate safely and effectively alongside humans in complex environments such as construction sites. While LiDAR-based SLAM and occupancy grid methods offer viable solutions for detecting and tracking motion, many state-of-the-art 3D vision approaches rely heavily on pre-trained neural networks and require additional post-processing to identify moving objects.

arXiv Computer Vision
Sep 18

AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend

AMB3R‑SLAM is a real‑time monocular SLAM system that can reconstruct kilometer‑scale trajectories over 10,000 frames on a single consumer‑grade GPU. It combines a lightweight front‑end for low‑latency tracking with a hierarchical backend that enforces local, mid‑level, and global consistency, avoiding bundle adjustment and thus handling dynamic scenes naturally. The system also supports stereo, RGB‑D, and LiDAR inputs, achieving strong camera tracking performance and reducing absolute trajectory error by over 70% on several datasets, with sub‑meter accuracy when LiDAR is added.

By Hengyi Wang, Lourdes Agapito
arXiv Computer Vision
Sep 23

Leveraging Vision-Based Point Cloud Map Priors for Camera-Based 3D Object Detection and Online Vectorized HD Mapping

The paper presents a framework that builds a static point cloud prior map from past camera traversals, augmenting each point with DINOv3 semantic features. During runtime, a local prior patch is retrieved, encoded with a sparse voxel backbone, and fused with lifted multi‑view camera features in bird’s‑eye view. This fused representation is then used by sparse transformer heads to predict 3D objects and vectorized map elements, achieving improved performance on Argoverse 2 without requiring LiDAR for prior‑map construction or online inference.

By Markus K\"appeler, Rohit Mohan, Abhinav Valada