arXiv Computer Vision By Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang

Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

Read the original on arXiv Computer Vision →

arXiv:2608. 20212v1 Announce Type: new Abstract: High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv Computer Vision
1d ago

Seeing Through the Glare: A Multi-Source Benchmark and Ocular-Adaptive Pixel MeanFlow for Eyeglass Reflection Removal

The paper introduces OcuBench, a comprehensive benchmark for eyeglass reflection removal that includes 10,280 synthetic pairs, 732 real-input pseudo-pairs, and 458 real-world test images, enabling both paired evaluation and assessment beyond generated supervision. It also proposes OcuFlow, an ocular-adaptive pixel MeanFlow framework that uses geometry-adaptive representation and one-step pMF to focus on reflection-obscured ocular regions while preserving native-resolution details. Experiments show OcuFlow consistently outperforms baselines in reflection removal quality, ocular fidelity, and efficiency, achieving 67.32% of selections in a blind user study, six times the next-best share.

By Tao Liu, Youwei Pang, Kailai Zhou, Jiaming Zuo, Hanqi Liu, Wei Ji, Peng-Tao Jiang, Xiaofeng Liu, Weisi Lin, Xiaoqi Zhao
arXiv Computer Vision
Aug 24

Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving

The paper introduces Driving with DINO (DwD), a framework that uses Vision Foundation Module (VFM) features to bridge simulation and real-world domains for autonomous driving video generation. It addresses the consistency‑realism dilemma by projecting VFM features onto a principal subspace, dropping high‑frequency texture elements, and applying a Random Channel Tail Drop to preserve structural detail. Additional components— a learnable Spatial Alignment Module and a Causal Temporal Aggregator— enhance control precision, spatial alignment, and temporal stability, reducing motion blur and ensuring realistic, consistent outputs.

By Xuyang Chen, Conglang Zhang, Chuanheng Fu, Zihao Yang, Kaixuan Zhou, Yizhi Zhang, Yanfeng Zhang, Mingwei Sun, Zhen Dong, Xiaoxiao Long, Zengmao Wang, Liqiu Meng