Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces ShadowCLR, an unsupervised framework for removing shadows from images without requiring paired shadow–shadow-free data or shadow masks. By leveraging consistency across multiple shadowed observations of the same scene, the method regularizes the model to recover scene-consistent appearance while suppressing shadow-specific variations. Experiments on several benchmarks show that ShadowCLR achieves competitive or superior performance compared to existing unsupervised approaches.
arXiv:2608. 06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems.
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision?
DiffusionShadow introduces a diffusion-based shadow caching framework for neural volume rendering, compressing many pre‑computed shadow INRs into a single diffusion model conditioned on lighting direction. The method encodes shadow coefficient volumes as shadow INRs, trains the diffusion model to predict shadow INR weights at inference, and integrates directly with standard INR renderers without extra runtime sampling. Experiments demonstrate faster rendering than traditional approaches while avoiding the large storage overhead of independent INRs, producing shadows that closely match reference results.
WildShadowRemover is a framework that adapts a pretrained video diffusion model for robust in-the-wild video shadow removal using LoRA fine-tuning. It augments the frozen VAE decoder with a detail injection module and introduces a shadow‑mask‑guided frequency‑decomposed modulation module to restore high‑frequency textures while suppressing shadow artifacts, with monocular depth priors providing geometry‑aware guidance. The authors also create WildShadow, a large‑scale paired video shadow removal dataset, and show that their method outperforms existing approaches in shadow removal quality, temporal consistency, and generalization across challenging real‑world scenarios.
arXiv:2609.00901v1 Announce Type: new Abstract: Modifying the illumination of driving images is a fundamental challenge, as most datasets are captured at specific times of day. Existing methods rely...