arXiv:2608. 06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems.
By Shilin Hu, Jingyi Xu, Dimitris Samaras, Hieu Le
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision?
The paper introduces ShadowCLR, an unsupervised framework for removing shadows from images without requiring paired shadow–shadow-free data or shadow masks. By leveraging consistency across multiple shadowed observations of the same scene, the method regularizes the model to recover scene-consistent appearance while suppressing shadow-specific variations. Experiments on several benchmarks show that ShadowCLR achieves competitive or superior performance compared to existing unsupervised approaches.
By Anh-Kiet Duong, Petra Gomez-Kr\"amer, Jean-Michel Carozza
WildShadowRemover is a framework that adapts a pretrained video diffusion model for robust in-the-wild video shadow removal using LoRA fine-tuning. It augments the frozen VAE decoder with a detail injection module and introduces a shadow‑mask‑guided frequency‑decomposed modulation module to restore high‑frequency textures while suppressing shadow artifacts, with monocular depth priors providing geometry‑aware guidance. The authors also create WildShadow, a large‑scale paired video shadow removal dataset, and show that their method outperforms existing approaches in shadow removal quality, temporal consistency, and generalization across challenging real‑world scenarios.
By Jiamin Xu, Cong Wang, Zheng Dong, Chi Wang, Renshu Gu, Weiwei Xu, Gang Xu
arXiv:2604.19254v2 Announce Type: replace
Abstract: Popular low-rank parameter-efficient fine-tuning (PEFT) methods represent adaptation as separate updates to selected backbone weights, without main...
By Xianming Li, Zongxi Li, Tsz-fung Andrew Lee, Jing Li, Haoran Xie, Qing Li
EviRover is a perception agent that goes beyond a single glance by actively gathering information to resolve perceptual queries. The authors created two data generation pipelines, producing EviRover-SFT-5K and EviRover-RL-12K, and a human‑verified benchmark called EviLens with 688 instances across five perception categories. Trained with supervised fine‑tuning and agentic reinforcement learning, the 4B EviRover outperforms its backbone by an average of 30 points on EviLens and shows strong transfer to other benchmarks such as WebEyes and BrowseComp‑VL.
By Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue