arXiv Computer Vision

A paired synthetic construction-site image dataset for robust computer vision under adverse conditions

arXiv Computer Vision
Sep 7

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

AdaptVPR introduces a route-aware generative augmentation framework that creates hard positive examples for Visual Place Recognition (VPR) training. It uses a vision‑language model to assess scene editability, a rule‑based scheduler to select generation routes, and a VPR‑oriented verification scheme to ensure geometric consistency and appearance diversity. The resulting AdaptCities dataset contains 160K verified synthetic hard positives, leading to consistent performance gains across VPR baselines, including up to 9.2% improvement in R@1 under challenging domain shifts.

By Shunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu
arXiv Computer Vision
Sep 7

Weather-Conditioned Depth Anything

Weather-Conditioned Depth Anything (DA‑W) is a new framework that enhances monocular depth estimation models, like the Depth Anything series, to perform robustly under adverse weather conditions such as fog, rain, snow, and low‑light. It achieves this by disentangling style from content: a Style Filter extracts weather‑specific embeddings from a curated mix of real and synthetic degradation data, which are then injected into the backbone via a lightweight, zero‑initialized adapter. The adapter is trained with pseudo‑label distillation and alignment, enabling a single unified model to adapt to diverse weather scenarios while preserving its generalization on clean data, and it achieves state‑of‑the‑art performance with an average 3.7% improvement in AbsRel on weather benchmarks.

By Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang, Zihao Zhu, Renjie Li, Yang Zhou, Zhengzhong Tu
arXiv AI
Jul 28

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.

By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
Hugging Face Trending Papers
Jun 23

Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity.

arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita