arXiv Computer Vision

Affine-Aligned Atlas for Canonical Gaussian Construction in Video Representation

The paper introduces an affine‑aligned atlas for constructing canonical Gaussian representations in video. By applying frame‑wise affine transforms to absorb global motion before building the canonical Gaussians, the method reduces misalignment between the canonical model and individual frames. This modification can be added to existing Gaussian‑based video representations with minimal extra parameters and improves reconstruction quality, particularly for sequences with large camera motion.

arXiv Computer Vision
Sep 25

GeoBlur: Epipolar Geometry Estimation from a Single Motion-Blurred Image

GeoBlur is a framework that estimates the fundamental matrix and relative camera pose from a single motion‑blurred image by exploiting blur artifacts as motion cues. It predicts visual correspondences between two time instances within the exposure window and solves the single‑frame epipolar geometry problem, yielding a fundamental matrix unique up to transposition due to time‑direction ambiguity. The method shows improved performance on synthetic and hybrid benchmarks and remains competitive on real motion‑blur data, also enabling downstream single‑frame motion segmentation.

By Bao-Long Tran, Cuong Le, Fredrik Viksten, Per-Erik Forss\'en
arXiv AI
Sep 28

OC-GS: Gaussian Splatting for Irregular Turntable Capture

OC-GS introduces an object‑centric Gaussian splatting method that refines the angle of each image while keeping a shared camera, rotation axis, and pivot for turntable reconstruction. By jointly optimizing image‑derived geometry and angles, it reconstructs objects from sparse, irregular captures and outperforms four pose‑free Gaussian splatting baselines across various view counts. Ablation studies confirm that both image‑derived angle initialization and the shared motion model are essential for the observed improvements, with real captures showing a 0.70 dB PSNR gain.

By Jae Joong Lee, Bedrich Benes
arXiv Computer Vision
Sep 18

VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors

VGGT-GS SLAM is a monocular 3D Gaussian Splatting SLAM system that operates on uncalibrated videos. It uses feed‑forward VGGT pose and depth priors to perform submap differentiable bundle adjustment, jointly refining camera poses, a 3D Gaussian map, and submap‑shared intrinsics and distortion parameters via analytic calibration Jacobians. The method introduces Gaussian‑native alignment for camera‑anchored scale refinement between submaps and loop‑closure verification, achieving improved localization accuracy and rendering quality on standard indoor benchmarks.

By Yuhang Han, Hao Wang, Jiaxi Cao, Xingyu Liu
arXiv Computer Vision
Sep 21

VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

VoxelTTO is a feed‑forward framework that reconstructs 3D Gaussian splatting scenes from multiple images by aggregating dense image features into a global voxel representation and decoding Gaussians from voxel features, thereby eliminating the pixel‑to‑Gaussian correspondence. It incorporates test‑time optimization with lightweight LoRA modules to adapt to known camera parameters while keeping the pretrained visual foundation model frozen. The method replaces standard rasterization with stochastic solid volume rendering, improving geometric fidelity, and demonstrates superior RGB‑D novel‑view synthesis and camera‑pose estimation on Replica, Tanks and Temples, and DTU datasets.

By Yibin Zhao, Yihan Pan, Yangwen Li, Jun Nan, Jianjun Yi
arXiv Computer Vision
Aug 27

Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-view Videos

Forge4D is a feed‑forward model that reconstructs temporally aligned 4D human representations from uncalibrated sparse‑view videos, enabling both novel view and novel time synthesis. It achieves this by jointly streaming 3D Gaussian reconstruction with dense motion prediction, using learnable state tokens for temporal consistency and a self‑supervised retargeting loss for motion prediction. Extensive experiments confirm its effectiveness on in‑domain and out‑of‑domain datasets.

By Yingdong Hu, Yisheng He, Jinnan Chen, Weihao Yuan, Kejie Qiu, Zehong Lin, Siyu Zhu, Zilong Dong, Steven Hoi, Jun Zhang