arXiv AI
Aug 24

Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

The paper proposes a structured method for assessing the representativeness of Operational Design Domains (ODDs) in AI/ML-based aviation systems, focusing on safety assurance. It outlines a process flow from ODD definition to quantitative evaluation, recommending Kullback–Leibler divergence and Cramér’s V over chi‑squared tests for large datasets. The approach is illustrated with AI-based collision avoidance simulations, demonstrating how statistical distribution comparisons can support safety‑by‑design engineering aligned with EASA guidance.

By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv AI
Sep 3

From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems

The paper proposes a structured method for verifying Operational Design Domain (ODD) coverage in safety‑critical AI systems, particularly for aviation. It combines parameter discretization, constraint‑based filtering, and criticality‑based dimension reduction to create a multi‑step verification process. Using simulation data from AI‑based mid‑air collision avoidance research, the authors demonstrate how this approach can meet EASA’s requirement for complete ODD coverage in high‑dimensional spaces.

By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach
arXiv AI
2d ago

Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

The Ontology-Based Contextual AI Evaluations (OB-CAIE) methodology introduces a structured approach to AI evaluation by defining clear testing coverage and balancing human expertise with automation. It employs two ontologies—the Domain‑Specific Ontology (DSO) outlining what is tested, and the Evaluation Process Ontology (EPO) detailing how it is tested—to create a tractable problem space that can be applied to single or multiple AI evaluations. OB‑CAIE enables traceable, visualizable failure points and incorporates human judgment at scientifically grounded junctures where machine input is insufficient.

By Julie Krugler Hollek, Michael Zargham, Mala Kumar