arXiv AI

Forward-Facing Near-Infrared Adds Little to Colour for Farm-Machinery Traversability: A Site-Disjoint Evaluation of Sensor-Dependent Spatial Leakage

arXiv Computer Vision
Sep 3

RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains

The paper explores using generative models to translate RGB UAV images into synthetic infrared (IR) images for training vehicle detectors in domains where real IR data is scarce. Various translators—supervised GANs, ControlNet-based diffusion models, and LoRA-ed foundation models—were trained on paired RGB-IR datasets and applied to unseen target datasets to generate synthetic IR data. The synthetic IR images, especially those produced by Stable Diffusion 3.5 with ControlNet, significantly improved detection performance on unseen IR test sets, outperforming RGB and grayscale baselines and narrowing the gap to real IR data.

By Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden, Elfi I. S. Hofmeijer, Sebastiaan P. Snel, Klamer Schutte, Friso G. Heslinga
arXiv AI
Jun 26

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

arXiv:2606. 26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge.

By Th\'eo Biardeau (XLIM-ASALI, UFR SFA), Anne-Sophie Capelle-Laiz\'e (UP, XLIM-ASALI, XLIM-ASALI), Salwan Alwan (UFR SFA), David Helbert (UFR SFA)
arXiv Computer Vision
Sep 23

Real-World Perception for Autonomous Driving in Adverse Weather: Enhancing Standard Detectors via Foundation-Guided Auto-Annotation

The paper presents a foundation-guided auto‑annotation pipeline that improves standard autonomous driving object detectors in adverse weather. By benchmarking YOLOv8, Co‑DETR, and SAM3 on a custom dataset of 25 operational scenarios, the authors find SAM3 to be the most robust and use it offline to generate pseudo‑labels. Fine‑tuning YOLOv8 on these labels boosts overall mAP by 16.04% and yields significant gains in specific conditions such as Residential Direct Sunlight (32.73%) and Highway Fog (28.65%).

By Sepideh Gohari, Goodarz Mehr, Azim Eskandarian
arXiv Computer Vision
Sep 17

Stealthy in Semantics, Antagonistic in Space: Attacking Visible-Infrared Object Detectors via Object-Level Misalignment

arXiv:2609.18133v1 Announce Type: new Abstract: Visible-infrared object detectors are used for robust perception under challenging illumination and weather conditions. Current physical attacks apply...

By Yueqi Zhu, Qi Ming, Guo Cheng, Yongkang Zhang, Feiran Liu, Juan Fang, Jiahuan Zhou, Jiangmeng Li, Yuhan Zhang
arXiv AI
Sep 3

InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

InfraPatch is a white‑box, per‑instance framework that generates small grayscale patches to target infrared‑adapted vision‑language models (IR‑VLMs). The method optimizes a single‑channel patch within a 5% local‑area budget, using proxy‑guided placement and task‑adaptive objectives to induce desired behaviors in image classification, captioning, and binary visual question answering. Across ten IR‑VLM variants tested on synthetic infrared images, InfraPatch achieves targeted attack success rates ranging from 86% to 100%, revealing significant vulnerability differences among architectures and tasks.

By Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo
arXiv AI
Aug 12

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

arXiv:2608. 11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming.

By Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
arXiv Machine Learning
Aug 31

Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

The paper introduces a depth‑aware pothole detection framework that fuses RGB‑D sensor data and evaluates five architectures—YOLOv8n, YOLOv8nSeg, YOLOv9t, RTDETRL, and RTDETRX—on the PothRGBD dataset. YOLOv8nSeg achieves the highest detection performance (mAP@50 = 0.9556, mAP@50_95 = 0.6758) and the most accurate depth estimate (2.96 cm), while YOLOv8n offers the fastest inference (3.6 ms) and RTDETRX delivers the highest detection confidence (92.70 %). The study also shows that even after RANSAC orthorectification, bounding‑box models overestimate pothole depth by 0.16–0.21 cm, indicating a structural bias rather than a calibration error.

By Md Monjurul Ahsan Prodhan, Md Nour Hossain