arXiv AI

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

arXiv:2603. 29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment.

arXiv AI
Sep 10

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.

By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita
Hugging Face Trending Papers
Jul 6

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.

arXiv AI
Jul 21

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

arXiv:2511. 14592v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns.

By Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao
arXiv AI
Jul 28

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.

By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
arXiv AI
Aug 25

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

arXiv:2608.21928v1 Announce Type: new Abstract: In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied...

By Zhesheng Zhang, Jiahao Lu, Wei Liu, Cong Pan, Jianhua Yang, Yixiang Chen, Hongyuan Yu, Mengqi Zhang, Kailin Lyu, Zhumin Chen, Keji He
arXiv AI
Sep 1

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

SafeAtlas-VL introduces a large multimodal safety dataset with 1.5 million instances, rating image, request, and response risks on a five‑level ordinal scale across 15 harm categories and 55 subcategories. The accompanying SafeAtlas‑Bench provides 5,000 held‑out cases for evaluating ordinal predictions and continuous risk scores. Models trained on this data, including an 8B Guard model, achieve state‑of‑the‑art performance, outperforming prior benchmarks by about 4% in F1 score.

By Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg