arXiv AI

Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety

arXiv:2606. 05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantified interventional effects, named root-cause variables), yet the XAI literature is organised by output type and technique family (saliency maps, feature attribution, counterfactuals, causal graphs, language traces).

arXiv AI
3d ago

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.

By Serhii Zabolotnii
arXiv AI
2d ago

A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging

The paper demonstrates that a verifier used in closed‑loop agent debugging can inadvertently reveal the answer it is meant to test, rendering solver comparisons meaningless. In a study of 12 development cases, both exact minimum hitting set and a greedy method returned identical supports, and an audit showed that exact‑anchor predicates always produced the planted fault pair. The authors propose a support‑gated verification contract that requires a clean reference map and runtime evidence before an independently calibrated signal can confirm a detection, and validate this approach on 1,440 held‑out cases with a low false‑admission rate.

By Peiying Zhu, Sidi Chang
arXiv Machine Learning
Sep 10

A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model

The paper presents a closed‑form estimator for the contamination correlation between anchors and judges under a single‑common‑factor model, requiring at least two judges and two anchors. It introduces a diagnostic battery—including judge‑covariance dispersion, over‑identification tests, a family‑block test, bootstrap confidence intervals, and a weak‑identification screen—to validate the estimator and detect violations. The authors also discuss identification limits for ordinal data and report that real panels have not yet passed the model‑adequacy pre‑test, while simulation studies confirm the estimator’s performance.

By Veerendra Kumar Sunkavalli
arXiv AI
2d ago

Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents

The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.

By Ali Atiah Alzahrani