The paper proposes a structured method for assessing the representativeness of Operational Design Domains (ODDs) in AI/ML-based aviation systems, focusing on safety assurance. It outlines a process flow from ODD definition to quantitative evaluation, recommending Kullback–Leibler divergence and Cramér’s V over chi‑squared tests for large datasets. The approach is illustrated with AI-based collision avoidance simulations, demonstrating how statistical distribution comparisons can support safety‑by‑design engineering aligned with EASA guidance.
By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv:2601.22118v3 Announce Type: replace
Abstract: Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems in...
By Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv:2608. 08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain.
By Chaitanya Shinde, Hadi Hajieghrary, Miguel Hurtado
arXiv:2608. 20053v1 Announce Type: new Abstract: The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment.
By Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann, Frank K\"oster, Sven Hallerbach
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-guided aggregation into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol revealed that safety compliance is the most discriminative dimension among 66 LLMs, with models showing up to 28-point differences in safety scores and recurrent failures such as safety violations under physically plausible predictions and instability in multi-step rollouts.
By Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
arXiv:2609.13552v1 Announce Type: new
Abstract: Generative AI is increasingly being used informally in Air Traffic Management (ATM) for tasks such as flight plan generation, trajectory interpretation...
By Alexandre Barreto (George Mason University), Shou Matsumoto (George Mason University), Jorge Valverde-Rebaza (Tecnol\'ogico de Monterrey), Cleiton Ataide (DECEA: Department of Airspace Control), Paulo Costa (George Mason University)
arXiv:2609.13642v1 Announce Type: new
Abstract: We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory complian...
By Hung-Yu Lin, Xingran Huang, Qiming Guo, Jinwen Tang
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints, then aggregates results into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol reveals that safety compliance is the most discriminative metric, with models of similar predictive accuracy differing by over 28 points in safety score and exhibiting recurrent safety violations and instability in multi-step rollouts.
arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
arXiv:2606. 17577v1 Announce Type: new Abstract: AI-driven engineering workflows face particular challenges in crash safety design: unlike aerodynamics, crash events involve highly nonlinear contact dynamics, material nonlinearity, and discrete state transitions that are difficult to capture with data-driven surrogate models.
By Osamu Ito, Akihiko Katagiri, Yoshikazu Nakagawa, Shin Saeki, Jun Shiraishi, Masato Sasaki
arXiv:2608.24621v2 Announce Type: replace
Abstract: Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee oper...
By Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam
arXiv:2511. 01650v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoning capabilities imperative.
By Ayesha Gull, Muhammad Usman Safder, Rania Elbadry, Fan Zhang, Veselin Stoyanov, Preslav Nakov, Zhuohan Xie