Hugging Face Trending Papers

Wat3R: Underwater 3D Geometry Learning without Annotations

Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings.

Hugging Face Trending Papers
Jul 23

WAT3R: Feedforward Underwater 3D Reconstruction

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate multi-view geometry. To address this issue, we propose WAT3R, a feed-forward framework for reconstructing 3D scenes directly from underwater images.

arXiv Computer Vision
Aug 31

3D-USE: From Image-Level to Scene-Level Underwater Enhancement

The paper introduces 3D-USE, a two‑stage framework for underwater scene‑level enhancement that learns a persistent, visibility‑enhanced 3D representation from degraded multi‑view observations. First, the Medium Radial Basis Anchor Representation (MediumRBF) builds a medium‑aware Gaussian scene by separating object and medium effects. Then, Appearance Transition Consensus (ATC) transfers 2D underwater image enhancement knowledge into scene‑global and Gaussian‑local targets, which are realized by an Underwater Bilateral Appearance Field (U‑BAF) to render enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes demonstrate improved visibility, cross‑view consistency, and preserved reconstruction quality.

By Jieyu Yuan, Yuanlin Zhang, Jihong Li, Chunle Guo, Huimin Lu, Chongyi Li
arXiv Computer Vision
Sep 25

OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning

OceanXL is a new framework that applies 3D Gaussian Splatting to large-scale underwater scenes by partitioning them into spatially coherent blocks and using adaptive pruning to remove redundant primitives. This divide‑and‑conquer approach improves training efficiency and rendering performance while maintaining global geometric consistency. The authors also release a large underwater dataset and demonstrate that OceanXL achieves favorable scalability, compactness, and efficiency compared to existing baselines, with competitive quality on smaller datasets and smaller model sizes than other underwater methods.

By Haoran Wang, Shaoyu Cai, Adrian Azzarelli, Zhuodong Jiang, Guoxi Huang, Eng Tat Khoo, Brett Seymour, Fan Zhang, David Bull, Nantheera Anantrasirichai
arXiv Computer Vision
Sep 25

WaterClear-GS: Optical-Aware Gaussian Splatting for Underwater Reconstruction and Restoration

WaterClear-GS introduces a physics-informed Gaussian splatting method tailored for underwater 3D reconstruction and appearance restoration. It models underwater degradation as intrinsic Gaussian attributes and employs a dual-branch optimization that separates clean appearance from degradation while preserving photometric consistency. The approach incorporates depth-guided geometry regularization, perception-driven supervision, exposure constraints, adaptive regularization, and spectral regularization, achieving strong novel view synthesis and image restoration performance at over 160 FPS.

By Xinrui Zhang, Yufeng Wang, Zesheng Wang, Dacheng Qi, Wenrui Ding, Shuangkang Fang
arXiv Computer Vision
Sep 18

Towards Scaling Marine Perception with Synthetic Data

The paper introduces an extension to the OceanSim underwater perception simulator, adding a Synthetic Data Generation pipeline that produces large, automatically labeled, photorealistic datasets with configurable scene and sensor settings. The authors evaluate this pipeline on a real-world sea urchin detection task, examining how different synthetic scene variations influence sim-to-real performance. They discuss the pipeline’s findings, limitations, and future directions for improving rendering fidelity, scene diversity, and sim-to-real generalization.

By Haoyu Ma, Onur Bagoren, Anja Sheppard, Elias Fandi, Ashrith Edukulla, Tanner Aslan, Natasha Sieh, Jingyu Song, Katherine A. Skinner
arXiv Computer Vision
Sep 7

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.

By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin