arXiv:2607. 18665v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance.
By Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding.
arXiv:2608. 12025v1 Announce Type: cross Abstract: Medical devices are becoming more software-intensive, connected, and AI-enabled.
By Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and r...
arXiv:2608. 04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level.
By Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo
arXiv:2604. 02022v4 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses.
By Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu
LogiScope‑VQA is a new benchmark dataset for evaluating vision‑language models in logistics hazard identification. It contains 2,476 images, 2,918 videos, and 10,274 VQA pairs drawn from real industrial warehouses, covering 18 core objects and 20 risk types across 39 subtasks. Experiments show that even advanced proprietary models lag behind human experts, highlighting a significant gap in perception, understanding, and reasoning for industrial safety.
By Hanjing Zhou, Mingze Yin, Ying Lian, Jun Ma, Chang-Yu Hsieh, Yanbing Zhou
The paper introduces DGEval, a benchmark of 1,678 questions designed to assess large language models (LLMs) on the International Maritime Dangerous Goods (IMDG) Code Amendment 42‑24. It evaluates 13 models from six providers, finding that while the best model surpasses human practitioners on multiple‑choice tasks, all models perform poorly on safety‑critical areas such as stowage, segregation, and regulatory recall. The study concludes that LLMs can aid compliance tasks—especially structured Dangerous Goods List lookups with web search—but human oversight and authoritative source verification remain essential for safety‑critical deployment.
By Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson
arXiv:2606. 14327v1 Announce Type: cross Abstract: This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspective of safety assurance.
By Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin
arXiv:2608.24621v2 Announce Type: replace
Abstract: Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee oper...
By Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam
The paper introduces a 202-scenario benchmark to evaluate how large language models (LLMs) handle safety-critical authorization decisions for vehicle voice commands. It tests two local open-weight models and three API-based LLMs, finding alignment scores ranging from 40.1% to 89.1% and noting persistent false execution errors. The study concludes that structured LLM decisions alone are insufficient for safety, recommending an independent enforcement layer to verify tool permissions and vehicle-state constraints before any vehicle function is invoked.
By Diba Afroze, Xingli Zhang, Yazhou Tu, Xiali Hei
arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.
By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang