arXiv:2606. 29181v1 Announce Type: cross Abstract: Detecting and localizing defects in 3D point clouds is challenging because abnormal samples are scarce and diverse, while training is often limited to normal data.
By Ali Balapour, Faraz Hach
Detecting and localizing defects in 3D point clouds is challenging because abnormal samples are scarce and diverse, while training is often limited to normal data. We propose Anomaly Factory 3D (AF3AD), a modular framework that synthesizes diverse pseudo-anomalies from normal point clouds to expand the training data for unsupervised 3D anomaly detection methods that rely on pseudo-anomalies.
AT3D-AD introduces a unified framework for detecting, localizing, and classifying 3D point‑cloud anomalies. It uses a Physics‑Driven Parametric Anomaly Synthesis module to generate synthetic defects for explicit supervision, a Hierarchical Global‑Local Anomaly Alignment module to refine representations, and a Semantic‑Geometric Anomaly Classification module to achieve precise, type‑discriminative localization. The method sets new state‑of‑the‑art results on four benchmarks, achieving high AUROC and Macro‑F1 scores.
By Jingyu Zeng, Haoquan Lu, Can Gao
PointGauss is a 3D-native framework that performs semantic parsing and instance segmentation on 3D Gaussian splatting representations by treating Gaussian primitives as unstructured point sets and extracting scale‑invariant geometric features with Point Transformer V3. It introduces an adaptive region‑of‑interest cropping strategy and an instance‑aware distance‑constrained rasterization pipeline to enable scalable, view‑consistent pixel‑level projections. The authors also release SplatSeg‑360, a cross‑scale benchmark with 32 complex scenes and over 6,300 aligned 2D‑3D masks, and show that PointGauss achieves real‑time performance with state‑of‑the‑art 3D‑mIoU (~90%) and 2D‑mIoU (~80%) scores.
By Wentao Sun, Yiping Chen, John S. Zelek, Jonathan Li
arXiv:2608. 07579v1 Announce Type: cross Abstract: The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only.
By Abdullah Naeem, Anav Katwal, Ayon Dey, Noman Khan, Md Tamjidul Hoque
arXiv:2404.09431v3 Announce Type: replace
Abstract: Pseudo-LiDAR has become a promising paradigm for monocular 3D object detection by transforming monocular images into point cloud representations th...
By Bonan Ding, Jin Xie, Jing Nie, Jiale Cao, Yanwei Pang
The paper addresses performance discrepancy in cross-domain 3D class‑incremental learning, where 3D point clouds from heterogeneous sources cause varying degrees of performance loss beyond catastrophic forgetting. The authors introduce the Domain3D‑CIL protocol and adapt existing CIL methods to 3D, showing consistent discrepancy across baselines. They propose PolyMem, an exemplar‑free approach that models high‑order feature statistics to improve cross‑domain robustness and reduce performance discrepancy.
By Jinge Ma, Gautham Vinod, Bruce Coburn, Jui-Feng Chi, Siddeshwar Raghavan, Fengqing Zhu
The paper challenges the common assumption that RGB and point cloud data contribute equally to zero‑shot multimodal anomaly detection. It shows that point clouds are more reliable under category shift and introduces WOOPS, a framework that enhances point cloud features with a Multi‑view Information Decoupling module and calibrates modality contributions via a Modality Reliability Calibration module. Experiments demonstrate that WOOPS achieves top performance on new stringent metrics in both unimodal and multimodal settings, and that point cloud information also benefits RGB‑only inference.
By Chenglin Ye, Lupeng Liu, Dongbo Yu, Jun Xiao, Yunbiao Wang
In modern high-throughput industrial production lines, product configurations and visual characteristics frequently change, making it impractical to collect and annotate data for every new scenario. This dynamic setting makes Zero-Shot Anomaly Detection (ZSAD) particularly suitable, as it enables defect detection without requiring training on target-specific samples.
arXiv:2608.30618v1 Announce Type: new
Abstract: Transformer-based decoders for 3D instance segmentation typically commit to a fixed number of queries and positional modeling calibrated on the trainin...
By Keno Moenck, Thorsten Sch\"uppstuhl
GAPrompt++ is a multi-granular geometry-aware prompting method designed to adapt pre-trained 3D vision models to downstream tasks efficiently. It introduces a Point Shift Prompter for multi-scale geometric feature extraction, a Keypoint Prompter for local geometric saliency, and a Prompt Propagation mechanism to embed these cues throughout the model hierarchy. Experiments demonstrate that GAPrompt++ outperforms other prompting-based PEFT methods and even surpasses full fine-tuning while using less than 2% trainable parameters, and the authors provide two new challenging benchmarks for future research.
By Zixiang Ai, Zhenyu Cui, Yufei Guo, Wenwen Qiang, Lei Chen, Jiwen Lu, Jiahuan Zhou
arXiv:2608. 07106v1 Announce Type: new Abstract: Deploying three-dimensional deep learning frameworks to low-power embedded processors is bottlenecked by the unstructured nature of spatial data and the resource-intensive distance sorting algorithms often used before neural network inference.
By Niclas Meyer, Stefan Reitmann