arXiv Computer Vision

Independent Samples, Correlated Variance A Learnable Cross-View Cue in Path-Traced Stereo Data

arXiv AI
Sep 2

VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM

VOIM (Voxel‑Grounded Online Instance Manager) is a training‑free system that builds open‑vocabulary 3D instance maps from RGB‑D or monocular RGB input by deferring label and instance decisions until sufficient soft evidence accumulates per voxel across views. Across four perception configurations on ScanNet++, VOIM outperforms the strongest online RGB‑D system, OVO‑SLAM, by 4.8–11.7 mIoU, and achieves 44.07 mIoU under a like‑for‑like protocol, winning all ten scenes. The method also runs unchanged on monocular RGB, matching baseline performance on Replica, and produces exportable occupancy grids that support free‑form instance queries.

By Sangmin Song, Sarath Kodagoda, Marc G. Carmichael, Karthick Thiyagarajan, Amal Gunatilake, Kelly Prentice, Jodi Martin
arXiv Computer Vision
Sep 4

Hold-Out Self-Validation Cannot Certify Photogrammetric Accuracy: Saturation and Blindness to Coherent Distortion

The paper argues that internal self-consistency checks cannot guarantee the accuracy of photogrammetric reconstructions, a limitation that is structural rather than a tuning issue. It introduces a track‑leakage‑free hold‑out protocol that withholds a deterministic subset of images and tests each against only 3D points supported by at least two retained images, ensuring no view is evaluated against the structure it helped create. Experiments on diverse datasets show that while the protocol is well‑posed, it saturates at a confidence score of 1.00 and fails to detect coherent distortion, missing large errors that can reach over 100 m. whyItMatters":"The study highlights that hold‑out self‑validation scores, increasingly used as quality evidence for metric deliverables, may be misleading and cannot replace external survey validation."

By Behnam Asadi
arXiv Computer Vision
4d ago

From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching

The paper introduces a 2D Gaussian Splatting pipeline that renders a dominant-eye RGB image and depth proxy, then reprojects and selectively patches the affiliated eye to reduce redundant work. By reusing alpha-blending weights and generating adaptive regions of interest, the method cuts sequential binocular rendering time by 15.5% to 28.8% and GPU memory by 6% to 11% on several datasets, with minimal quality loss. It demonstrates a practical efficiency‑quality trade‑off for static‑scene stereo rendering and suggests further evaluation on dynamic scenes and VR hardware.

By Hongfei Zhu, Ling Zhou
arXiv AI
Sep 3

Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging

The paper investigates why post‑hoc saliency maps, such as Grad‑CAM, shift when input images are rotated, even if the model’s prediction remains unchanged. By measuring equivariance at each stage of the CAM operator, the authors find that the instability originates from the spatial activation tensor rather than the channel weights, and that a training‑free wrapper called EquiGrad‑CAM can align and average saliency maps across rotated views to significantly improve rotation equivariance. Experiments on ImageNet, PatchCamelyon, and RESISC45 demonstrate that EquiGrad‑CAM outperforms rotation‑augmented training and enhances zero‑shot CLIP explanations.

By Khawaja Murad ul Hassan, Mehran Ebrahimi
arXiv AI
Aug 26

Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

Syn2RealTrack addresses the synthetic‑to‑real gap in multi‑camera 3D perception for warehouses by decomposing it into three distinct issues: camera calibration, object shape prior, and known object census. The pipeline corrects lens distortion from images, fuses detections with a visibility‑weighted part‑based descriptor, measures person height directly from calibration, and uses a closed‑world cardinality prior with a causal filter to eliminate phantom boxes. These local remedies allow the system to adapt without retraining a feature extractor, achieving a 3D HOTA of 52.0118% on the AI City Challenge 2026 Track 1.

By Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc Pham-Nam Ho, Trinh Le Ba Khanh, Chi Dai Tran, Duong Khac Vu, Son Hong Phan, Hyung-Min Jeon, Jae Wook Jeon
arXiv AI
Sep 10

CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

CALIPER is a new benchmark that tests whether pretrained visual encoders can infer physical properties such as mass and friction from images. The test involves striking an object twice at known speeds, showing a third strike only up to contact, and asking a linear readout on frozen features to predict how far the object slides. Results show that in clean, fixed‑camera scenes all representations perform similarly, but when camera, lighting, and clutter are varied, only encoders that truly infer physics—like V‑JEPA 2—maintain performance, while random or raw pixel representations fail.

By Aman Mehta, Riya Baviskar