GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.
By Yufei Zhang, Chenlu Zhan, Hongwei Wang
arXiv:2605.30320v2 Announce Type: replace
Abstract: Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D...
By Daniel Rho, Jun Myeong Choi, Matthew Thornton, Biswadip Dey, Roni Sengupta
arXiv:2608.31025v1 Announce Type: new
Abstract: Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remain...
By Jailing Lin, Jikuan Zhang, Jianhua Sun
The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
By Jung-Hee Kim, Xiaoming Liu
SymRegFlow is a symmetry‑regularized flow‑matching framework that enables multi‑view‑consistent video generation across continuously varying camera poses without requiring ground‑truth novel‑view RGB supervision. The method geometrically warps source views into noisy anchors and uses masked dual‑anchor supervision combined with cross‑anchor denoising‑output consistency to reduce anchor‑specific errors. Experiments on Cosmos‑Drive‑Dreams and nuScenes show that SymRegFlow achieves superior video quality, achieving the lowest FVD and FVMD scores and improving FID and instance preservation compared to existing baselines.
By Xi Ye, Yuzhu Wang, Xiaoyang Liu, Jiayi Wang, Yangyang Xu, Ruyu Wang, Wenlin Chen, Duo Su, Jun Zhu
GS‑VLA introduces a lightweight, plug‑and‑play framework that uses a 4 M‑parameter 3D‑Gaussian canonicalizer to adapt frozen Vision‑Language‑Action (VLA) policies to viewpoint shifts without retraining the policy. By treating viewpoint changes as a localized novel‑view synthesis problem under a locality assumption, the method normalizes observations through a scene‑ and policy‑independent disocclusion task. Experiments on the LIBERO benchmark demonstrate that GS‑VLA recovers a large portion of performance lost due to camera displacement, improving results across different policy architectures, unseen task suites, and perturbation scales.
whyItMatters":"The approach offers a computationally efficient alternative to costly fine‑tuning or generative augmentation, enabling robust VLA deployment in real‑world settings where camera configurations may vary."
By Yechan Park, HyunJin Kim
arXiv:2609.36929v1 Announce Type: new
Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...
By Thai Duy Nguyen, Addison Lin Wang
arXiv:2604.09045v2 Announce Type: replace
Abstract: Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level...
By Tsuheng Hsu, Guiyu Liu, Juho Kannala, Janne Heikkil\"a
The paper introduces GRF-Recon, a framework for stable and scalable feed-forward 3D reconstruction from long monocular image sequences. It combines coarse-to-fine trajectory alignment, lightweight geometric prior injection via LoRA adaptation, and a hybrid-weight sparse ray-field optimization to refine local point clouds while enforcing cross-frame consistency. An efficient trajectory stitching strategy with joint ray-error optimization further reduces accumulated drift, achieving competitive trajectory accuracy compared to SLAM systems while maintaining globally consistent reconstructions in large-scale scenarios.
By Enpeng Li, Yunzhou Zhang, Zhiyao Zhang, Dexuan Lyu, Chenyu Wang, Chiyuan Cui, Cheng Cheng
The paper introduces 3D Morphological Perturbations, an optimization‑free regularizer for 3D representations such as NeRF and 3D Gaussian Splatting. By treating each Gaussian as a pixel‑like element, the method applies scale, rotation, and pruning perturbations to preserve spatial consistency across views, eliminating the need for per‑scene optimization during dataset curation. Experiments on a lightweight video diffusion sandbox and a 14B‑parameter video model show that the approach improves geometric priors, reduces mean depth error by 12.5% over state‑of‑the‑art 3D artifact refiners, and boosts downstream robotics policy success rates by up to 8.0% on three manipulation tasks.
By Onat \c{S}ahin, Mohammad Altillawi, George Eskandar, Carlos Carbone, Ziyuan Liu
arXiv:2608.22740v1 Announce Type: new
Abstract: Generalizable 3D Gaussian Splatting (G-3DGS) has emerged as a promising approach for novel view synthesis undersparse-view settings. However, existing...
By Zeyang Bai, Yunpeng Wang, Yunbiao Wang, Jun Xiao
MoSE3 is a feed‑forward model that predicts dense SE(3) motion—full 6‑DoF rigid transforms—at every pixel from monocular RGB video, providing rotation, translation, and grouping information simultaneously. It achieves this by jointly learning 3D point tracks and rigidity embeddings, then differentiably fitting transforms within soft rigid clusters. The authors also release Art‑Kubric, a large synthetic dataset with dense SE(3) and rigidity labels, and demonstrate state‑of‑the‑art performance on both rigid and articulated benchmarks, with strong generalization to real‑world videos.
By Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang