arXiv Computer Vision

Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction

The paper introduces a state‑aware framework that reconstructs both foreground and hidden scenes from single‑photon LiDAR histograms affected by partially transmissive occluders. It classifies each ray into no‑return, single‑return, or dual‑return states, guiding a two‑head neural field that jointly refines waveform reconstruction and geometry localization. A new paired dataset of occluded and clean LiDAR captures validates the method, showing improved depth and point‑cloud accuracy over existing baselines.

Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.

arXiv Computer Vision
Sep 25

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.

By Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno
arXiv Computer Vision
Sep 1

OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes

arXiv:2608.29881v1 Announce Type: new Abstract: Monocular depth estimation has achieved strong open-domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective...

By Muxin Liu, Tianbo Liu, Jing Xia, Xiaoyang Lyu, Xiaoshan Wu, Bo Wang, Peng Dai, Zhongrui Wang, Shaoshuai Shi, Xiaojuan Qi
arXiv Computer Vision
Sep 18

Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks

The paper presents a lightweight 3D U‑Net designed to detect ultra‑sparse LiDAR occupancy of bat flight paths in nocturnal field recordings. By preserving temporal resolution and combining weighted binary cross‑entropy with Dice loss, the model overcomes the class imbalance that hampers standard reconstruction methods. Experiments on real LiDAR data, cross‑checked with acoustic monitoring, show that the U‑Net successfully recovers coherent occupancy patterns along bat trajectories, offering a practical foundation for large‑scale validation, clustering of flight tracks, and integration into biodiversity‑aware turbine curtailment strategies.

By Nico Klar, Pankaj Rana, Nizam Gifary, Jakob Traub, Aamir Ahmad
arXiv Computer Vision
Sep 1

RLG-TPV: Radar- and LiDAR-Guided Tri-Perspective View Fusion for Camera-Radar 3D Object Detection

RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.

By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates