arXiv Computer Vision

LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering

LensStyle is a unified framework for controllable stylized lens effect rendering that explicitly models lens aesthetics through joint continuous‑discrete control. It uses a Dual‑Path Controller to separate continuous optical parameter modulation (e.g., focus distance, blur strength) from discrete lens‑style conditioning (e.g., circular, polygonal, donut, cat‑eye, starburst effects), allowing fine‑grained, interpretable, and physically grounded manipulation. The authors also curate a MultiLens dataset of multi‑lens image pairs synthesized under real optical constraints, and experiments show LensStyle outperforms existing lens effect rendering methods and diffusion‑based image editing models in realism, controllability, and aesthetic quality.

arXiv Computer Vision
Aug 21

Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

arXiv:2608. 20212v1 Announce Type: new Abstract: High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections.

By Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang
Hugging Face Trending Papers
Jul 7

Realistic Compound-Lens Defocus Blur Synthesis

Defocus blur degrades fine image structures and limits visual perception, which can adversely affect downstream vision tasks. Although recent deep learning deblurring methods have achieved strong performance, their effectiveness depends on training data and often degrades across cameras and lenses due to limited optical diversity and realism in existing datasets.

arXiv AI
Sep 10

Aes3D: Aesthetic Assessment in 3D Gaussian Splatting

arXiv:2605.05155v4 Announce Type: replace-cross Abstract: As 3D Gaussian Splatting (3DGS) gains attention in immersive media and digital content creation, assessing the aesthetics of 3D scenes become...

By Chuanzhi Xu, Boyu Wei, Haoxian Zhou, Xuanhua Yin, Zihan Deng, Haodong Chen, Qiang Qu, Weidong Cai
arXiv Machine Learning
Jun 3

Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors

arXiv:2511. 17126v4 Announce Type: replace-cross Abstract: Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by training a universal neural network, demonstrating strong capability in handling diverse unknown optical degradations.

By Xiaolong Qian, Qi Jiang, Yao Gao, Lei Sun, Kailun Yang, Xian Wang, Zhonghua Yi, Wenyong Li, Ming-Hsuan Yang, Luc Van Gool, Kaiwei Wang
arXiv AI
Jun 30

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

arXiv:2512. 17504v2 Announce Type: replace-cross Abstract: Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D scene understanding and a lack of proper optical interactions, such as shadows and reflections.

By Hoiyeong Jin, Hyojin Jang, Junha Hyung, Jeongho Kim, Kinam Kim, Dongjin Kim, Huijin Choi, Hyeonji Kim, Jaegul Choo
arXiv Computer Vision
Sep 18

GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets

GS-PI introduces an optimization‑decoupled framework that transforms Gaussian Splatting (GS) assets into physically based rendering (PBR) compatible Gaussian assets. By treating PBR material generation as a geometry‑conditioned diffusion process on 3D point clouds, it achieves multi‑view consistency and avoids the pixel‑correspondence problems of 2D diffusion. The method employs a multi‑scale cross‑view conditioning mechanism—combining global semantic priors, photometric cues, and spatial view‑direction signals—to prevent specular highlights from baking into intrinsic colors, and then distills the predicted attributes back into a fully relightable PBR‑GS asset without requiring proxy meshes.

By Jieting Xu, Rengan Xie, Zijian Huang, Zehui Jin, Rui Wang, Yuchi Huo
arXiv Computer Vision
Sep 2

JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

JanusMesh introduces a fast, training‑free framework for creating 3D visual illusion meshes that reveal different semantics from various viewpoints. The method splits generation into two stages: a cross‑space dual‑branch denoising process that aligns 3D latents with CLIP guidance and blends Signed Distance Fields for seamless geometry, followed by a view‑conditioned texture synthesis module that aggregates 2D diffusion priors onto the fused mesh. Experiments show that JanusMesh produces highly realistic, dual‑semantic 3D illustrations in only 3–5 minutes, outperforming prior approaches in geometric integrity, semantic recognizability, and efficiency.

By Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang, Yu-Lun Liu
arXiv Computer Vision
Sep 7

Collaborative On-Sensor Array Cameras

The paper presents a collaborative on‑sensor array camera that uses a distributed meta‑optics learning method to jointly optimize a 100‑million‑nanopost metasurface array for broadband visible imaging. By training the array end‑to‑end with a learned meta‑atom proxy and a parallax‑aware, noise‑aware reconstruction algorithm, the design overcomes the wavelength‑dependent limitations of traditional metalenses. Experimental results show that the camera delivers consistent image quality across varying scene illumination spectra without relying on generative reconstruction.

By Jipeng Sun, Kaixuan Wei, Thomas Eboli, Congli Wang, Cheng Zheng, Zhihao Zhou, Arka Majumdar, Wolfgang Heidrich, Felix Heide