Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Grounded-Exo2Ego introduces a dual‑branch video diffusion model that combines a geometric anchoring branch with a semantic grounding branch to generate egocentric video from a single exocentric source. The framework also includes a camera re‑localization algorithm to correct reconstruction misalignment and a fully automated synthetic data engine for training. Experiments on the EgoExo4D dataset demonstrate significant performance gains over recent state‑of‑the‑art methods.
arXiv:2609.38615v1 Announce Type: cross Abstract: Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Ex...
arXiv:2606. 02753v1 Announce Type: cross Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
LIFT is a unified image‑to‑video generation framework that adds Layout‑In‑Future control, letting users specify what should appear and where in a future view. It addresses the limitation of existing camera controls and text prompts by using the last‑frame layout as an explicit signal for the desired future scene, especially under large viewpoint changes. To handle sparse layout guidance, LIFT employs on‑policy self‑distillation to transfer knowledge from a dense‑layout teacher to a last‑frame‑layout student, and introduces the LIFT‑Vista dataset with large viewpoint changes and consistent layout annotations. Experiments demonstrate that LIFT improves video quality, future‑layout controllability, and camera controllability compared to other methods.
arXiv:2607. 17790v1 Announce Type: cross Abstract: Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment.
arXiv:2606. 31585v1 Announce Type: cross Abstract: The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is crucial for providing spatial cues in multi-view geometry.