arXiv AI
Sep 2

VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM

VOIM (Voxel‑Grounded Online Instance Manager) is a training‑free system that builds open‑vocabulary 3D instance maps from RGB‑D or monocular RGB input by deferring label and instance decisions until sufficient soft evidence accumulates per voxel across views. Across four perception configurations on ScanNet++, VOIM outperforms the strongest online RGB‑D system, OVO‑SLAM, by 4.8–11.7 mIoU, and achieves 44.07 mIoU under a like‑for‑like protocol, winning all ten scenes. The method also runs unchanged on monocular RGB, matching baseline performance on Replica, and produces exportable occupancy grids that support free‑form instance queries.

By Sangmin Song, Sarath Kodagoda, Marc G. Carmichael, Karthick Thiyagarajan, Amal Gunatilake, Kelly Prentice, Jodi Martin
arXiv Computer Vision
Sep 4

Hold-Out Self-Validation Cannot Certify Photogrammetric Accuracy: Saturation and Blindness to Coherent Distortion

The paper argues that internal self-consistency checks cannot guarantee the accuracy of photogrammetric reconstructions, a limitation that is structural rather than a tuning issue. It introduces a track‑leakage‑free hold‑out protocol that withholds a deterministic subset of images and tests each against only 3D points supported by at least two retained images, ensuring no view is evaluated against the structure it helped create. Experiments on diverse datasets show that while the protocol is well‑posed, it saturates at a confidence score of 1.00 and fails to detect coherent distortion, missing large errors that can reach over 100 m. whyItMatters":"The study highlights that hold‑out self‑validation scores, increasingly used as quality evidence for metric deliverables, may be misleading and cannot replace external survey validation."

By Behnam Asadi