arXiv Computer Vision

LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image

The paper introduces LLPR, a framework that integrates location-aware learning and physics-based reconstruction to remove raindrops from single images. It replaces costly preprocessing masks with a learnable branch that can be discarded during inference, and reconstructs the background by first estimating a transparency matrix and raindrop layer using a physical model. The authors also present a new real-world dataset and show that LLPR outperforms existing state‑of‑the‑art methods.

arXiv Computer Vision
Sep 7

Weather-Conditioned Depth Anything

Weather-Conditioned Depth Anything (DA‑W) is a new framework that enhances monocular depth estimation models, like the Depth Anything series, to perform robustly under adverse weather conditions such as fog, rain, snow, and low‑light. It achieves this by disentangling style from content: a Style Filter extracts weather‑specific embeddings from a curated mix of real and synthetic degradation data, which are then injected into the backbone via a lightweight, zero‑initialized adapter. The adapter is trained with pseudo‑label distillation and alignment, enabling a single unified model to adapt to diverse weather scenarios while preserving its generalization on clean data, and it achieves state‑of‑the‑art performance with an average 3.7% improvement in AbsRel on weather benchmarks.

By Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang, Zihao Zhu, Renjie Li, Yang Zhou, Zhengzhong Tu
arXiv Machine Learning
Sep 10

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

arXiv:2609.08084v1 Announce Type: cross Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...

By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
arXiv Computer Vision
Sep 7

ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization

ARC‑Loc introduces a new cross‑view localization method that bypasses heavy Bird’s‑Eye‑View transformations and external depth models. By converting ground keypoints into azimuthal rays on a satellite map and exploiting their convergence at the user’s location, the approach uses a minimal Azimuthal Ray Convergence solver and an ARC loss to directly match ground and satellite images. Experiments on VIGOR and KITTI show that ARC‑Loc achieves competitive accuracy while offering faster, memory‑efficient inference and easy integration with existing frameworks.

By Hyeongsik Kim, Mincheol Kim, Heejoon Moon, Je Hyeong Hong
arXiv AI
Jul 7

MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model

arXiv:2607. 03013v1 Announce Type: cross Abstract: Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines, which degrades visual quality and affects downstream vision tasks.

By Wanshu Fan, Xiangyu Li, Cong Wang, Kin-man Lam, Xin Yang, Haiyan Zhang, Dongsheng Zhou