WildShadowRemover is a framework that adapts a pretrained video diffusion model for robust in-the-wild video shadow removal using LoRA fine-tuning. It augments the frozen VAE decoder with a detail injection module and introduces a shadow‑mask‑guided frequency‑decomposed modulation module to restore high‑frequency textures while suppressing shadow artifacts, with monocular depth priors providing geometry‑aware guidance. The authors also create WildShadow, a large‑scale paired video shadow removal dataset, and show that their method outperforms existing approaches in shadow removal quality, temporal consistency, and generalization across challenging real‑world scenarios.
By Jiamin Xu, Cong Wang, Zheng Dong, Chi Wang, Renshu Gu, Weiwei Xu, Gang Xu
arXiv:2610.01291v1 Announce Type: new
Abstract: Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world...
By Junseong Shin, Kijun Kim, Minseong Kim, Dongjin Kim, Tae Hyun Kim
The paper introduces ShadowCLR, an unsupervised framework for removing shadows from images without requiring paired shadow–shadow-free data or shadow masks. By leveraging consistency across multiple shadowed observations of the same scene, the method regularizes the model to recover scene-consistent appearance while suppressing shadow-specific variations. Experiments on several benchmarks show that ShadowCLR achieves competitive or superior performance compared to existing unsupervised approaches.
By Anh-Kiet Duong, Petra Gomez-Kr\"amer, Jean-Michel Carozza
MAST (Mask‑Guided Attention Control for Training‑Free Regional‑Multi Style Transfer) is a framework that enables diffusion models to apply multiple reference styles to user‑specified regions of a content image without any training or optimization. It introduces logit‑level attention mass allocation, sharpness‑aware temperature scaling, and discrepancy‑aware detail injection to address mass allocation, selectivity, and detail loss problems in regional‑multi style transfer. Experiments with two to five styles show that MAST outperforms baselines in ArtFID, FID, and R‑FID, achieving high regional style fidelity, content preservation, and scalability.
By Dongkyung Kang, Jaeyeon Hwang, Junseo Park, Minji Kang, Yeryeong Lee, Beomseok Ko, Hanyoung Roh, Jeongmin Shin, Hyeryung Jang
arXiv:2609.01123v1 Announce Type: new
Abstract: Recent advancements in low-light image enhancement have leveraged diffusion models for their strong ability to generate perceptually realistic, detaile...
By Ruoyu Guo, Haonan Zhong, Maurice Pagnucco, Yang Song
arXiv:2605. 31162v1 Announce Type: cross Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored.
By Shreyansh Modi, Akshat Tomar, Aarush Aggarwal
arXiv:2608.29243v1 Announce Type: new
Abstract: Existing diffusion-based enhancement methods provide strong generative capability for low-light image enhancement (LLIE), yet they either rely on paire...
By Wenjie Cai, Yuezhe Yang, Jianyang Xia, Xingbo Dong, Zhe Jin
arXiv:2606. 04299v1 Announce Type: cross Abstract: We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that of a single reference image.
By Haojun Qiu, Kiriakos N. Kutulakos, David B. Lindell
DiffusionShadow introduces a diffusion-based shadow caching framework for neural volume rendering, compressing many pre‑computed shadow INRs into a single diffusion model conditioned on lighting direction. The method encodes shadow coefficient volumes as shadow INRs, trains the diffusion model to predict shadow INR weights at inference, and integrates directly with standard INR renderers without extra runtime sampling. Experiments demonstrate faster rendering than traditional approaches while avoiding the large storage overhead of independent INRs, producing shadows that closely match reference results.
By Kai-Chen Tung, Qi Wu, David Bauer, Mengjiao Han, Silvio Rizzi, Kwan-Liu Ma
Recent advancements in low-light image enhancement have leveraged diffusion models for their strong ability to generate perceptually realistic, detailed images. Patch diffusion models further offer a...
RelightFormer is a feed‑forward generative Transformer that performs single‑ and multi‑view image relighting without explicit intrinsic property estimation. It incorporates a latent illumination module that injects target environment maps into spatial features via cross‑attention, and uses permutation‑invariant positional encodings to process unordered multi‑view inputs symmetrically. Trained on the large Laval Objaverse Dataset, the model achieves state‑of‑the‑art visual and photorealistic relighting quality, and demonstrates strong zero‑shot generalization across various relighting tasks.
By Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
arXiv:2601. 11641v3 Announce Type: replace-cross Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment.
By Yuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang, Kun Yuan