LogiScope‑VQA is a new benchmark dataset for evaluating vision‑language models in logistics hazard identification. It contains 2,476 images, 2,918 videos, and 10,274 VQA pairs drawn from real industrial warehouses, covering 18 core objects and 20 risk types across 39 subtasks. Experiments show that even advanced proprietary models lag behind human experts, highlighting a significant gap in perception, understanding, and reasoning for industrial safety.
By Hanjing Zhou, Mingze Yin, Ying Lian, Jun Ma, Chang-Yu Hsieh, Yanbing Zhou
SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability.
"whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."
By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv:2603. 04818v3 Announce Type: replace Abstract: Disruptions at critical logistics nodes pose severe risks to global supply chains, yet existing risk prediction systems typically prioritize forecasting accuracy without providing operationally interpretable early warnings.
By Zhiming Xue, Yujue Wang, Menghao Huo
arXiv:2607. 16243v1 Announce Type: cross Abstract: Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying.
By Anushiya Arunan, Xin Li, Yan Qin, U-Xuan Tan, Nhu Khue Vuong, Xiaoli Li, Chau Yuen
arXiv:2608. 09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment.
By Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren
arXiv:2608. 11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views.
By Xikai Sun, Cangtian Zhou, Kebin Liu, Ke Ma, Xu Wang, Zaishu Chen, Haotian Wang, Li Liu, Yunhao Liu
Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU).
arXiv:2607. 06625v1 Announce Type: cross Abstract: Process industries have accumulated validated specialist models, yet sensor drift, feedstock variation, and regime switching cause these models to degrade systematically in new scenarios.
By Youcheng Zong, Runda Jia, Ranmeng Lin, Mingxuan Ren, Dakuo He
arXiv:2607. 18325v1 Announce Type: cross Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies.
By Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu, Stephanie Lukin, Cynthia Matuszek
The paper introduces a two‑stage training framework that combines Supervised Fine‑Tuning (SFT) and Direct Preference Optimization (DPO) to improve multimodal disaster severity assessment. It creates two datasets—ReasoningSet for validated rationales and PreferenceSet for paired rationales—using a single Human‑in‑the‑Loop workflow. Experiments on InternVL‑3‑8B and LLaVA‑1.5‑7B show that SFT boosts classification accuracy and Macro‑F1, while DPO further enhances interpretability and alignment with human judgment.
By Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee, Fahad Khalid, Mourad Oussalah
The paper introduces DGEval, a benchmark of 1,678 questions designed to assess large language models (LLMs) on the International Maritime Dangerous Goods (IMDG) Code Amendment 42‑24. It evaluates 13 models from six providers, finding that while the best model surpasses human practitioners on multiple‑choice tasks, all models perform poorly on safety‑critical areas such as stowage, segregation, and regulatory recall. The study concludes that LLMs can aid compliance tasks—especially structured Dangerous Goods List lookups with web search—but human oversight and authoritative source verification remain essential for safety‑critical deployment.
By Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson
arXiv:2606. 22437v2 Announce Type: replace-cross Abstract: We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results.
By Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong, Chengping Zhao, Ting Liu, Yuzhuo Fu