arXiv AI

After a Decade: Bringing Shadow Removal into the Real World with Agentic Training Data

The paper introduces AgenticShadow, a new dataset of 17,138 image‑mask‑target triplets created through an offline agentic workflow that combines physics‑motivated generation, failure detection, feedback‑driven retry, candidate selection, and deterministic correction. This approach addresses the long‑standing lack of diverse paired shadow‑free training data by leveraging existing shadow detection datasets and producing realistic shadow‑free targets. Models trained on AgenticShadow show significant improvements, reducing color distribution differences by 50.5% and cross‑domain LAB RMSE by 19.7‑37.5% compared to prior work.

arXiv Computer Vision
Sep 3

Consistency as Regularization for Unsupervised Shadow Removal

The paper introduces ShadowCLR, an unsupervised framework for removing shadows from images without requiring paired shadow–shadow-free data or shadow masks. By leveraging consistency across multiple shadowed observations of the same scene, the method regularizes the model to recover scene-consistent appearance while suppressing shadow-specific variations. Experiments on several benchmarks show that ShadowCLR achieves competitive or superior performance compared to existing unsupervised approaches.

By Anh-Kiet Duong, Petra Gomez-Kr\"amer, Jean-Michel Carozza
arXiv Computer Vision
Aug 25

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

WildShadowRemover is a framework that adapts a pretrained video diffusion model for robust in-the-wild video shadow removal using LoRA fine-tuning. It augments the frozen VAE decoder with a detail injection module and introduces a shadow‑mask‑guided frequency‑decomposed modulation module to restore high‑frequency textures while suppressing shadow artifacts, with monocular depth priors providing geometry‑aware guidance. The authors also create WildShadow, a large‑scale paired video shadow removal dataset, and show that their method outperforms existing approaches in shadow removal quality, temporal consistency, and generalization across challenging real‑world scenarios.

By Jiamin Xu, Cong Wang, Zheng Dong, Chi Wang, Renshu Gu, Weiwei Xu, Gang Xu
arXiv AI
1d ago

EviRover: Reinforcing Agentic Perception Beyond a Glance

EviRover is a perception agent that goes beyond a single glance by actively gathering information to resolve perceptual queries. The authors created two data generation pipelines, producing EviRover-SFT-5K and EviRover-RL-12K, and a human‑verified benchmark called EviLens with 688 instances across five perception categories. Trained with supervised fine‑tuning and agentic reinforcement learning, the 4B EviRover outperforms its backbone by an average of 30 points on EviLens and shows strong transfer to other benchmarks such as WebEyes and BrowseComp‑VL.

By Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv AI
4d ago

Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement

Consist‑Retinex introduces a one‑step noise‑emphasized consistency training framework for Retinex‑based low‑light image enhancement. It first decomposes images into reflectance and illumination maps using a Retinex Transformer Decomposition Network, then trains two conditional consistency models with a dual objective that blends trajectory consistency and ground‑truth alignment. The method employs adaptive noise‑emphasized fixed‑point sampling to focus supervision near the inference endpoint, achieving state‑of‑the‑art VE‑LOL‑L scores on paired and unpaired low‑light benchmarks while reducing sampling and training costs.

By Jian Xu, Wei Chen, Shigui Li, Delu Zeng, John Paisley, Qibin Zhao
arXiv Computer Vision
Sep 7

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.

By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng