Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings.
arXiv:2608.22906v1 Announce Type: new
Abstract: Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality...
By Yingxiang Xu, Kerui Ren, Wenqi Guo, Changjian Jiang, Tao Lu, Linning Xu, Mulin Yu
Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality and efficiency. However, extending these framewo...
OceanXL is a new framework that applies 3D Gaussian Splatting to large-scale underwater scenes by partitioning them into spatially coherent blocks and using adaptive pruning to remove redundant primitives. This divide‑and‑conquer approach improves training efficiency and rendering performance while maintaining global geometric consistency. The authors also release a large underwater dataset and demonstrate that OceanXL achieves favorable scalability, compactness, and efficiency compared to existing baselines, with competitive quality on smaller datasets and smaller model sizes than other underwater methods.
By Haoran Wang, Shaoyu Cai, Adrian Azzarelli, Zhuodong Jiang, Guoxi Huang, Eng Tat Khoo, Brett Seymour, Fan Zhang, David Bull, Nantheera Anantrasirichai
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward vi...
WaterClear-GS introduces a physics-informed Gaussian splatting method tailored for underwater 3D reconstruction and appearance restoration. It models underwater degradation as intrinsic Gaussian attributes and employs a dual-branch optimization that separates clean appearance from degradation while preserving photometric consistency. The approach incorporates depth-guided geometry regularization, perception-driven supervision, exposure constraints, adaptive regularization, and spectral regularization, achieving strong novel view synthesis and image restoration performance at over 160 FPS.
By Xinrui Zhang, Yufeng Wang, Zesheng Wang, Dacheng Qi, Wenrui Ding, Shuangkang Fang
arXiv:2508. 04928v5 Announce Type: replace-cross Abstract: We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images.
By Rit Gangopadhyay, Jung-Hee Kim, Xien Chen, Patrick Rim, Hyoungseob Park, Alex Wong
The paper introduces 3D-USE, a two‑stage framework for underwater scene‑level enhancement that learns a persistent, visibility‑enhanced 3D representation from degraded multi‑view observations. First, the Medium Radial Basis Anchor Representation (MediumRBF) builds a medium‑aware Gaussian scene by separating object and medium effects. Then, Appearance Transition Consensus (ATC) transfers 2D underwater image enhancement knowledge into scene‑global and Gaussian‑local targets, which are realized by an Underwater Bilateral Appearance Field (U‑BAF) to render enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes demonstrate improved visibility, cross‑view consistency, and preserved reconstruction quality.
By Jieyu Yuan, Yuanlin Zhang, Jihong Li, Chunle Guo, Huimin Lu, Chongyi Li
AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.
By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often depend on heavy decoder stacks and lack a structured mechanism for geometry refinement, resulting in poor multi-view consistency.
arXiv:2608.22888v1 Announce Type: new
Abstract: Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpr...
By Xiaopeng Guo, Wai Chung Tse, Yipeng Zhu, Hanwen Zhang, Huajian Huang, Sai-Kit Yeung
The paper introduces Distortion Extenders (DEX), learnable parameters that adapt vision foundation models to fisheye cameras by modeling distortion coefficients and correcting distributional shifts between fisheye and perspective images. DEX is applied to monocular depth estimation and open‑vocabulary segmentation across convolutional and Transformer architectures, consistently outperforming baselines on indoor and outdoor fisheye datasets. Additionally, DEX activations can be decoded to obtain distortion coefficients, aiding camera calibration.
By Rit Gangopadhyay, Alex Wong