Hugging Face Trending Papers

WAT3R: Feedforward Underwater 3D Reconstruction

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate multi-view geometry. To address this issue, we propose WAT3R, a feed-forward framework for reconstructing 3D scenes directly from underwater images.

arXiv Computer Vision
Sep 25

OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning

OceanXL is a new framework that applies 3D Gaussian Splatting to large-scale underwater scenes by partitioning them into spatially coherent blocks and using adaptive pruning to remove redundant primitives. This divide‑and‑conquer approach improves training efficiency and rendering performance while maintaining global geometric consistency. The authors also release a large underwater dataset and demonstrate that OceanXL achieves favorable scalability, compactness, and efficiency compared to existing baselines, with competitive quality on smaller datasets and smaller model sizes than other underwater methods.

By Haoran Wang, Shaoyu Cai, Adrian Azzarelli, Zhuodong Jiang, Guoxi Huang, Eng Tat Khoo, Brett Seymour, Fan Zhang, David Bull, Nantheera Anantrasirichai
arXiv Computer Vision
Sep 25

WaterClear-GS: Optical-Aware Gaussian Splatting for Underwater Reconstruction and Restoration

WaterClear-GS introduces a physics-informed Gaussian splatting method tailored for underwater 3D reconstruction and appearance restoration. It models underwater degradation as intrinsic Gaussian attributes and employs a dual-branch optimization that separates clean appearance from degradation while preserving photometric consistency. The approach incorporates depth-guided geometry regularization, perception-driven supervision, exposure constraints, adaptive regularization, and spectral regularization, achieving strong novel view synthesis and image restoration performance at over 160 FPS.

By Xinrui Zhang, Yufeng Wang, Zesheng Wang, Dacheng Qi, Wenrui Ding, Shuangkang Fang
arXiv Computer Vision
Aug 31

3D-USE: From Image-Level to Scene-Level Underwater Enhancement

The paper introduces 3D-USE, a two‑stage framework for underwater scene‑level enhancement that learns a persistent, visibility‑enhanced 3D representation from degraded multi‑view observations. First, the Medium Radial Basis Anchor Representation (MediumRBF) builds a medium‑aware Gaussian scene by separating object and medium effects. Then, Appearance Transition Consensus (ATC) transfers 2D underwater image enhancement knowledge into scene‑global and Gaussian‑local targets, which are realized by an Underwater Bilateral Appearance Field (U‑BAF) to render enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes demonstrate improved visibility, cross‑view consistency, and preserved reconstruction quality.

By Jieyu Yuan, Yuanlin Zhang, Jihong Li, Chunle Guo, Huimin Lu, Chongyi Li
arXiv Computer Vision
Sep 7

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.

By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
Hugging Face Trending Papers
Jun 2

BA-T: An Iterative Transformer for Two-View Bundle Adjustment

Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often depend on heavy decoder stacks and lack a structured mechanism for geometry refinement, resulting in poor multi-view consistency.

arXiv Computer Vision
Aug 31

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

The paper introduces Distortion Extenders (DEX), learnable parameters that adapt vision foundation models to fisheye cameras by modeling distortion coefficients and correcting distributional shifts between fisheye and perspective images. DEX is applied to monocular depth estimation and open‑vocabulary segmentation across convolutional and Transformer architectures, consistently outperforming baselines on indoor and outdoor fisheye datasets. Additionally, DEX activations can be decoded to obtain distortion coefficients, aiding camera calibration.

By Rit Gangopadhyay, Alex Wong