Future Dynamic 3D Reconstruction: Toward 3D World Modeling with Disentangled Ego-Motion
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2603. 03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities.
SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.
GenIA is a framework that aligns generative 3D foundation models with test‑time observations, improving pose estimation and reconstruction from monocular, multi‑view, and dynamic inputs. It derives translation and scale from geometry, retains learned rotation priors, and aligns appearance using visibility‑biased attention, cross‑observation fusion, and differentiable rendering guidance during denoising. An optional post‑denoising refinement further adapts appearance latents and object placement, and the method also supports externally supplied geometry for dynamic objects, achieving better results than recent optimization‑based and image‑to‑3D methods on synthetic and real benchmarks.
Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.
Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.
arXiv:2607. 17097v2 Announce Type: replace Abstract: Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI.