Hugging Face Trending Papers

CUBICS: Situation-aware performance estimation for safety-relevant ML components

Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.

arXiv Machine Learning
Sep 17

Uncertainty measurement for complex event prediction in safety-critical systems

The paper presents a machine‑learning approach (ML_CP) that automatically learns patterns and rules for complex event prediction, reducing reliance on manual rule creation. It incorporates sensitivity analysis to assess how output varies with each input and uses conformal prediction to generate uncertainty‑aware prediction intervals. Experiments on binary, multi‑level classification, and regression tasks show promising results for safety‑critical embedded systems.

By Maria J. P. Peixoto, Akramul Azim
arXiv Machine Learning
Jun 19

Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence

arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.

By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv Computation and Language
Sep 18

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability. "whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."

By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv AI
Aug 24

Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

The paper proposes a structured method for assessing the representativeness of Operational Design Domains (ODDs) in AI/ML-based aviation systems, focusing on safety assurance. It outlines a process flow from ODD definition to quantitative evaluation, recommending Kullback–Leibler divergence and Cramér’s V over chi‑squared tests for large datasets. The approach is illustrated with AI-based collision avoidance simulations, demonstrating how statistical distribution comparisons can support safety‑by‑design engineering aligned with EASA guidance.

By Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach