arXiv Computer Vision
Aug 24

Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Grounded-Exo2Ego introduces a dual‑branch video diffusion model that combines a geometric anchoring branch with a semantic grounding branch to generate egocentric video from a single exocentric source. The framework also includes a camera re‑localization algorithm to correct reconstruction misalignment and a fully automated synthetic data engine for training. Experiments on the EgoExo4D dataset demonstrate significant performance gains over recent state‑of‑the‑art methods.

By Shengze Wang, Michael Stengel, Tianye Li, Seonwook Park, Amrita Mazumdar, Koki Nagano, Alex Trevithick, Shalini De Mello
arXiv Computer Vision
2d ago

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

LIFT is a unified image‑to‑video generation framework that adds Layout‑In‑Future control, letting users specify what should appear and where in a future view. It addresses the limitation of existing camera controls and text prompts by using the last‑frame layout as an explicit signal for the desired future scene, especially under large viewpoint changes. To handle sparse layout guidance, LIFT employs on‑policy self‑distillation to transfer knowledge from a dense‑layout teacher to a last‑frame‑layout student, and introduces the LIFT‑Vista dataset with large viewpoint changes and consistent layout annotations. Experiments demonstrate that LIFT improves video quality, future‑layout controllability, and camera controllability compared to other methods.

By Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu