CADSplat is a framework that reconstructs photorealistic, geometrically accurate digital twins from fewer than 15 wide‑baseline images by regularizing 3D Gaussian Splatting with an explicit CAD shape prior. It matches segmented object silhouettes to a CAD library to retrieve a suitable model and camera poses, then anchors Gaussian primitives to the model’s surface and jointly optimizes splat parameters, registration, and a non‑rigid deformation field. Experiments on two real‑world datasets show CADSplat outperforms baselines, especially in sparse and self‑occluded scenarios, and its gains mainly stem from constraining splats to a surface rather than the CAD shape itself.
By Kristof Overdulve, Lode Jorissen, Nick Michiels
The paper presents a method for tracking insects in the field using a stereoscopic event-based camera setup. By converting asynchronous events into conventional video formats, the authors combine the high temporal resolution of event cameras with standard video processing techniques to capture detailed insect flight movements. The stereoscopic configuration enables continuous, low‑latency 3D tracking, reducing motion blur and improving accuracy in natural environments.
By Pratham G. Shenwai, Martin J. Lankheet, John T. Hrynuk, Mandiyam Y. Mahadeeswara, Mandyam V. Srinivasan, Sridhar Ravi
arXiv:2609.10376v1 Announce Type: new
Abstract: Sparse-view X-ray 3D reconstruction is essential for reducing radiation exposure, but recovering a density field from a handful of X-ray projections is...
By Pranav Poudel, Florence Dell'Aniello Picard, Nairouz Shehata, Fr\'ed\'eric Lavoie, Herve Lombaert
arXiv:2609.24253v1 Announce Type: cross
Abstract: 3D Gaussian Splatting (3DGS) provides high-fidelity scenes for large-scale embodied simulation, but constructing large-scale urban assets remains con...
By Zhongrui You, Zhen Li, Junli Liu, Zhigang Wang, Bin Zhao
Sparse-view X-ray 3D reconstruction is essential for reducing radiation exposure, but recovering a density field from a handful of X-ray projections is severely ill-posed. Recently, 3D Gaussian Splatt...
arXiv:2605.30320v2 Announce Type: replace
Abstract: Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D...
By Daniel Rho, Jun Myeong Choi, Matthew Thornton, Biswadip Dey, Roni Sengupta
ORMA is a training‑free framework that reconstructs articulated 4D representations of animals from monocular videos by decoupling pose and shape. It uses predicted pose as a reference for optimization and generative 3D priors to refine shape, aligning the result with the SMAL+ parametric model. The method combines per‑frame pose estimates with globally consistent camera poses, and further refines the reconstruction using self‑supervised DINO correspondences and temporal consistency, achieving improved accuracy on the new PAW4D benchmark and diverse real‑world videos.
By Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers, Riccardo Marin, Silvia Zuffi
The paper introduces Camera Splatting, a novel framework for optimizing camera viewpoints in novel view synthesis. Each camera is represented as a 3D Gaussian (camera splat), and virtual point cameras are positioned near the surface to sample the distribution of these splats. By continuously refining the camera splats to match desired target distributions observed from the point cameras, the method achieves better capture of complex view‑dependent effects such as metallic reflections and detailed textures compared to the Farthest View Sampling approach.
By Gahye Lee, Hyomin Kim, Gwangjin Ju, Jooeun Son, Hyejeong Yoon, Seungyong Lee
MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.
By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang
Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.
By Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery.
arXiv:2608.28386v1 Announce Type: new
Abstract: Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic o...
By Semin Kim, Haechan Shin, Jongyoo Kim