The paper introduces Ref-GeNVS, a training‑free, reflection‑aware approach for generative novel view synthesis in mirror scenes. It treats a mirror image as two complementary views, estimates the mirror plane and reflected camera poses, and uses a two‑stage generation process with Mirror‑gated attention and Reflection injection to produce reflection‑consistent novel views. The method leverages a multi‑view diffusion backbone without finetuning, outperforming recent generative NVS methods on synthetic and real mirror scenes.
By GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh
The paper introduces a method for mirror inpainting that leverages scene geometry to generate realistic reflections. By estimating the geometry of the fixed scene, the approach projects visible content into the mirror region, reducing the need for hallucination. A two‑mask diffusion strategy then refines the mirror area, balancing geometric constraints with learned priors, and the method operates without training on complex real‑world scenes.
By Ofek Basson, Shimon Vainer, Yacov Hel-Or, Ohad Fried
arXiv:2609.23442v1 Announce Type: new
Abstract: Diffusion models generate high-quality images, yet often violate the physical laws governing mirror reflections. Reflections often suffer from geometri...
By Shuheng Ge, Hongwei Ren, Li Zhang, Xiangqian Wu
arXiv:2607.03470v2 Announce Type: replace
Abstract: Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasing...
By Xuan-Bach Mai, Duy-Phuc Nguyen, Quoc-Van Le, Tam V. Nguyen, Thanh-Toan Do, Huu Le, Duong-Van Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv:2608. 11562v1 Announce Type: cross Abstract: Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks.
By Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
arXiv:2609.35734v2 Announce Type: replace
Abstract: Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, whi...
By Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
arXiv:2506.01004v3 Announce Type: replace-cross
Abstract: Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybr...
By Tong Zhang, Victor Escorcia, Juan C Leon Alcazar, Bernard Ghanem
arXiv:2608. 07559v1 Announce Type: cross Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models.
By Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu
CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.
By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.
By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
ContextAnyone is a context‑aware diffusion framework that treats a reference image as an explicitly preserved appearance anchor rather than a simple conditioning signal. By jointly reconstructing the reference image and generating the target video within a shared diffusion transformer, it provides direct supervision for maintaining identity and fine‑grained appearance throughout denoising. The method introduces asymmetric information flow and Gap‑RoPE positional representations to keep the reference stable while allowing selective access by video tokens, and demonstrates improved identity and appearance consistency on an OpenVid‑HD benchmark.
By Ziyang Mai, Yu-Wing Tai