arXiv AI

Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction

Hugging Face Trending Papers
Aug 10

XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors' released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work.

arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
arXiv Machine Learning
Sep 25

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.

By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv Computer Vision
Sep 17

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

The paper introduces the Wide-area Spatio-temporal Scene Understanding (WSTU) problem, which demands simultaneous wide-area coverage, per-target resolution, and temporal continuity—capabilities lacking in existing datasets. To address this, the authors present HARD, an ultra‑high‑resolution (12768×9564) UAV dataset annotated for object detection, multi‑object tracking, and scene‑level visual question answering. They also propose a latency‑aware metric, streaming‑HOTA (s‑HOTA), and show through baseline experiments that high resolution and processing latency significantly impact detection, tracking, and VQA performance, revealing gaps in current methods for WSTU.

By Yuhang Zhu, Meiyi Zhu, Yunkai Dang, Zhangnan Li, Yuxuan Wang, Wenbin Li, Hongbing Pan