arXiv Computer Vision
Sep 28

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales introduces TangoGS, a method that automatically selects the number of Gaussian primitives for 3D Gaussian Splatting by combining capture-derived model sizing with training-based adaptation. The approach first estimates a learning allowance based on the capture’s total pixels, then adjusts the number of Gaussians during training according to reconstruction quality. On standard benchmarks, TangoGS matches the best baseline’s PSNR while using 48% fewer Gaussians, and on larger captures it scales automatically to achieve the highest mean PSNR with 2.3× more Gaussians.

By Afif Boudaoud, Jiayi Liu, Alexandru Calotoiu, Torsten Hoefler
arXiv Computer Vision
Sep 18

Open-vocabulary 3D object detection with promptable segmentation

The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.

By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
arXiv Computer Vision
Sep 14

Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation

The paper investigates how much semantic information is lost when frozen foundation models are combined for few‑shot 3D segmentation. By varying the number of retained semantic alternatives before fusion, the authors show that keeping the full distribution of class scores yields higher harmonic‑mean IoU than collapsing to a single class. Experiments on ScanNet200 and ScanNet++ confirm that full‑distribution fusion consistently outperforms top‑1 and other operators, and that most useful information is recovered by retaining a compact set of plausible alternatives.

By Silas Kwabla Gah, Ebenezer Owusu