arXiv AI By Faizan Iqbal

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

Read the original on arXiv AI →

arXiv:2607. 20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg