arXiv Computer Vision

Practical High-Fidelity Novel-View Synthesis of Mounted Lepidoptera

arXiv Computer Vision
Sep 17

CADSplat: Sparse-View 3D Gaussian Splatting Aided by CAD Models for Robust, Photorealistic Digital-Twin Reconstruction

CADSplat is a framework that reconstructs photorealistic, geometrically accurate digital twins from fewer than 15 wide‑baseline images by regularizing 3D Gaussian Splatting with an explicit CAD shape prior. It matches segmented object silhouettes to a CAD library to retrieve a suitable model and camera poses, then anchors Gaussian primitives to the model’s surface and jointly optimizes splat parameters, registration, and a non‑rigid deformation field. Experiments on two real‑world datasets show CADSplat outperforms baselines, especially in sparse and self‑occluded scenarios, and its gains mainly stem from constraining splats to a surface rather than the CAD shape itself.

By Kristof Overdulve, Lode Jorissen, Nick Michiels
arXiv Computer Vision
Sep 21

Field Tracking of Insects Using a Stereoscopic Event-Based Camera Setup

The paper presents a method for tracking insects in the field using a stereoscopic event-based camera setup. By converting asynchronous events into conventional video formats, the authors combine the high temporal resolution of event cameras with standard video processing techniques to capture detailed insect flight movements. The stereoscopic configuration enables continuous, low‑latency 3D tracking, reducing motion blur and improving accuracy in natural environments.

By Pratham G. Shenwai, Martin J. Lankheet, John T. Hrynuk, Mandiyam Y. Mahadeeswara, Mandyam V. Srinivasan, Sridhar Ravi
arXiv Computer Vision
6d ago

ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals

ORMA is a training‑free framework that reconstructs articulated 4D representations of animals from monocular videos by decoupling pose and shape. It uses predicted pose as a reference for optimization and generative 3D priors to refine shape, aligning the result with the SMAL+ parametric model. The method combines per‑frame pose estimates with globally consistent camera poses, and further refines the reconstruction using self‑supervised DINO correspondences and temporal consistency, achieving improved accuracy on the new PAW4D benchmark and diverse real‑world videos.

By Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers, Riccardo Marin, Silvia Zuffi
arXiv Computer Vision
Sep 4

Camera Splatting for Continuous View Optimization

The paper introduces Camera Splatting, a novel framework for optimizing camera viewpoints in novel view synthesis. Each camera is represented as a 3D Gaussian (camera splat), and virtual point cameras are positioned near the surface to sample the distribution of these splats. By continuously refining the camera splats to match desired target distributions observed from the point cameras, the method achieves better capture of complex view‑dependent effects such as metallic reflections and detailed textures compared to the Farthest View Sampling approach.

By Gahye Lee, Hyomin Kim, Gwangjin Ju, Jooeun Son, Hyejeong Yoon, Seungyong Lee
arXiv Computer Vision
4d ago

MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation

MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.

By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang
arXiv Computer Vision
Sep 4

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

By Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy