The paper proposes a structured method for verifying Operational Design Domain (ODD) coverage in safety‑critical AI systems, particularly for aviation. It combines parameter discretization, constraint‑based filtering, and criticality‑based dimension reduction to create a multi‑step verification process. Using simulation data from AI‑based mid‑air collision avoidance research, the authors demonstrate how this approach can meet EASA’s requirement for complete ODD coverage in high‑dimensional spaces.
By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
The paper proposes a structured method for assessing the representativeness of Operational Design Domains (ODDs) in AI/ML-based aviation systems, focusing on safety assurance. It outlines a process flow from ODD definition to quantitative evaluation, recommending Kullback–Leibler divergence and Cramér’s V over chi‑squared tests for large datasets. The approach is illustrated with AI-based collision avoidance simulations, demonstrating how statistical distribution comparisons can support safety‑by‑design engineering aligned with EASA guidance.
By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv:2608. 20053v1 Announce Type: new Abstract: The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment.
By Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv:2608. 04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data.
By Chinmoy Mitra, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Rakibul Islam, M. F. Mridha
Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data. This study examines two complementary challenges: benign heterogeneity, where honest operators observe different operating conditions and fault modes, and adversarial heterogeneity, where compromised operators submit poisoned updates.
arXiv:2410. 22526v2 Announce Type: replace Abstract: To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards.
By Shalaleh Rismani, Roel Dobbe, AJung Moon
arXiv:2609.13552v1 Announce Type: new
Abstract: Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation...
By Alexandre Barreto (George Mason University), Shou Matsumoto (George Mason University), Jorge Valverde-Rebaza (Tecnol\'ogico de Monterrey), Cleiton Ataide (DECEA: Department of Airspace Control), Paulo Costa (George Mason University)
Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-guided aggregation into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol revealed that safety compliance is the most discriminative dimension among 66 LLMs, with models showing up to 28-point differences in safety scores and recurrent failures such as safety violations under physically plausible predictions and instability in multi-step rollouts.
By Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
arXiv:2608. 16564v1 Announce Type: new Abstract: Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications.
By Benjamin Herd, Jessica Kelly, Mario Trapp
arXiv:2411.01289v2 Announce Type: replace
Abstract: Complex events originate from other primitive events combined according to defined patterns and rules. Instead of using specialists' manual work to...
By Maria J. P. Peixoto, Akramul Azim
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints, then aggregates results into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol reveals that safety compliance is the most discriminative metric, with models of similar predictive accuracy differing by over 28 points in safety score and exhibiting recurrent safety violations and instability in multi-step rollouts.