arXiv:2608. 16564v1 Announce Type: new Abstract: Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications.
By Benjamin Herd, Jessica Kelly, Mario Trapp
The paper presents a machine‑learning approach (ML_CP) that automatically learns patterns and rules for complex event prediction, reducing reliance on manual rule creation. It incorporates sensitivity analysis to assess how output varies with each input and uses conformal prediction to generate uncertainty‑aware prediction intervals. Experiments on binary, multi‑level classification, and regression tasks show promising results for safety‑critical embedded systems.
By Maria J. P. Peixoto, Akramul Azim
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.
By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv:2410. 22526v2 Announce Type: replace Abstract: To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards.
By Shalaleh Rismani, Roel Dobbe, AJung Moon
arXiv:2604.14251v2 Announce Type: replace
Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be es...
By Edoardo Pona, Milad Kazemi, Mehran Hosseini, Yali Du, David Watson, Osvaldo Simeone, Nicola Paoletti
arXiv:2608. 13793v1 Announce Type: cross Abstract: Machine learning (ML) has become an indispensable part of modern engineering design workflows.
By Tyler R. Johnson, Kian Ben-Jacob, Christopher P. Muller, Ramin Bostanabad
SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability.
"whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."
By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv:2609.16373v1 Announce Type: new
Abstract: Probabilistic certification of Bayesian neural networks lower-bounds the posterior probability that a model satisfies a verifier-defined safety propert...
By Mahyar Mohammadi, Mohammad Hossein Badiei, Abolfazl Yaghmaei, Hamed Kebriaei
The paper proposes a structured method for assessing the representativeness of Operational Design Domains (ODDs) in AI/ML-based aviation systems, focusing on safety assurance. It outlines a process flow from ODD definition to quantitative evaluation, recommending Kullback–Leibler divergence and Cramér’s V over chi‑squared tests for large datasets. The approach is illustrated with AI-based collision avoidance simulations, demonstrating how statistical distribution comparisons can support safety‑by‑design engineering aligned with EASA guidance.
By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
By Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli
arXiv:2602. 13562v2 Announce Type: replace-cross Abstract: While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures.
By Yanbo Wang, Minzheng Wang, Jian Liang, Lu Wang, Yongcan Yu, Ran He