arXiv Computer Vision

Geometry vs Structure: Graph-Based Diagnostics for LiDAR Point-Cloud Simulation Fidelity

arXiv Computer Vision
Sep 3

GEM: Generating LiDAR World Model via Deformable Mamba

GEM is a Generative LiDAR world model that uses a deformable Mamba architecture to better handle the disorder of LiDAR point clouds and distinguish dynamic objects from static structures. The model tokenizes LiDAR sweeps, unsupervisedly disentangles dynamic and static features, and applies a tri‑path deformable Mamba for selective scanning and adaptive gating fusion, improving spatial‑temporal understanding. Experiments show GEM outperforms existing methods across multiple benchmarks, and it can be paired with a planner and BEV controller for autonomous rollout and "what‑if" scenario generation.

By Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.

Hugging Face Trending Papers
Jul 23

DTIF: Robust Loop Closure Detection via Delaunay Triangle Topology in Complex Forests

Accurate forest inventory and large-scale mapping are essential for ecosystem monitoring and sustainable forest management. Multiple low-cost edge platforms enable efficient large-area data acquisition, but merging independently constructed local maps in GNSS-denied understory environments still requires initialization-free loop closure detection and global registration.

arXiv Computer Vision
Sep 21

PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection

PointLAM introduces a new point-based 3D object detection architecture that addresses efficiency and fidelity trade-offs inherent in LiDAR point cloud processing. It employs a Laplacian Point Sampler (LPS) to accelerate downsampling while preserving foreground structure, and a Local Hadamard Aggregator (LHA) that replaces costly continuous interactions with a topology‑aware gating mechanism. Combined with Bi‑Directional Mamba layers, the resulting Local Attentive Mamba (LAM) block delivers competitive performance on nuScenes and Waymo datasets, outperforming voxel‑based competitors in detecting small objects and handling extreme sparsity with a smaller computational footprint.

By Xuanming Shang, Weijia Zhang, Chao Ma
Hugging Face Trending Papers
Jun 1

Honey, I Shrunk the Arc de Triomphe!

Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collapse'' phenomenon: distant landmarks and vast landscapes are metrically underestimated. We hypothesize that this performance gap stems from a training data bottleneck, where existing metric-scale datasets are hardware-constrained to homogenous vehicle-captured LiDAR or short-range indoor scans, or consist of synthetic data that lacks the semantic complexity of the physical world.