arXiv AI

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

arXiv:2512. 17504v2 Announce Type: replace-cross Abstract: Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D scene understanding and a lack of proper optical interactions, such as shadows and reflections.

arXiv Computer Vision
2d ago

4Director: Controlling Video World Models with Rigid 3D Geometry

arXiv:2610.02160v1 Announce Type: new Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...

By Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
Hugging Face Trending Papers
Jul 27

CameraAnything: Refilming Videos with Arbitrary Camera Control

We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation.

arXiv Computer Vision
Sep 7

WorldSculpt: Generating Compositional Worlds from Grounded Videos

WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.

By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
arXiv Computer Vision
Aug 24

Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Grounded-Exo2Ego introduces a dual‑branch video diffusion model that combines a geometric anchoring branch with a semantic grounding branch to generate egocentric video from a single exocentric source. The framework also includes a camera re‑localization algorithm to correct reconstruction misalignment and a fully automated synthetic data engine for training. Experiments on the EgoExo4D dataset demonstrate significant performance gains over recent state‑of‑the‑art methods.

By Shengze Wang, Michael Stengel, Tianye Li, Seonwook Park, Amrita Mazumdar, Koki Nagano, Alex Trevithick, Shalini De Mello
arXiv Computer Vision
Aug 21

Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

arXiv:2608. 20212v1 Announce Type: new Abstract: High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections.

By Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang