Hugging Face Trending Papers

GRAR: Glass-induced Reflection Artifact Removal in LiDAR Point Clouds

Terrestrial Laser Scanning (TLS) point clouds captured in urban environments frequently suffer from glass-induced reflection artifacts, severely degrading downstream applications. Existing reflection artifact removal methods generally rely on ideal reflection symmetry assumptions, yet their performance is limited by inaccurate glass estimation and insufficient geometric representations.

arXiv Computer Vision
Aug 28

Glass Surface Detection Grounded in 3D Visual Geometry

Glass Surface Detection Grounded in 3D Visual Geometry proposes a new approach that grounds glass surface detection in 3D visual geometry rather than relying solely on 2D appearance cues. The method uses a visual geometry grounded transformer (VGGT) to distill 3D priors and creates glass-aware 3D representations, then applies a multi-task learning framework with a Frequency Self-Attention Module (FSAM) and a Geometry Grounding Block (GeGB) to localize and segment glass surfaces. Experiments show state‑of‑the‑art performance on seven benchmarks, good generalization to video and multi‑modal data, and significant improvements in reconstruction of glass scenes.

By Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau
arXiv Computer Vision
Aug 25

3D Point Cloud from Close-Range Photogrammetry for Defect Characterisation of Rubberised Concrete

The paper presents a close‑range photogrammetry workflow using Structure‑from‑Motion and Multi‑View Stereo to generate high‑resolution 3D point clouds of rubberised concrete. By capturing images with a Canon DSLR and an iPhone 16, the authors achieved sub‑millimetre reconstruction accuracy, outperforming traditional LiDAR for fine‑scale defect analysis. An RGB‑guided crack extraction method and deformation analysis further demonstrate the method’s utility for detailed surface monitoring and material performance evaluation.

By Jiacheng Liu, Mohammed Alnahhal, Ailar Hajimohammadi, Sara Gonizzi Barsanti, Jinling Wang, Mohsen Kalantari
Hugging Face Trending Papers
Aug 20

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations.

arXiv AI
Aug 21

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

arXiv:2608. 19536v1 Announce Type: cross Abstract: Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics.

By Eunsoo Im, Junghun Suh, Gyeonggwan Lee, Seunghwan Hong
arXiv Computer Vision
Sep 25

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.

By Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.