arXiv AI

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

arXiv:2608. 03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action.

arXiv Machine Learning
Sep 22

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

SafeTune is a source‑available library that consolidates four safety‑intervention paradigms—post‑hoc weight recovery, safety‑constrained fine‑tuning, gradient‑based unlearning, and inference‑time steering—into a single, configuration‑driven workflow. It offers shared interpretability, evaluation, and deployment tools, and its modular registry allows easy addition of new methods, benchmarks, judges, models, and fine‑tuning domains. The authors demonstrate SafeTune with controlled comparisons and case studies in finance and medical deployments, showing how it characterizes safety drift, evaluates interventions on refusal‑behavior and capability metrics, and supports calibrated or layered mitigation.

By Pratinav Seth, Saisab Sadhu, Anshul Kaushal, Vinay Kumar Sankarapu
arXiv AI
3d ago

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.

By Serhii Zabolotnii
arXiv Computation and Language
Sep 18

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability. "whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."

By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu