arXiv AI

ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models

arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody
arXiv Machine Learning
Jul 27

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.

By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
arXiv Machine Learning
Sep 11

SoK: Privacy Attacks on Machine Learning via Explainable AI

The paper surveys 25 studies that use explainable AI to compromise machine learning models, covering attacks such as model extraction, membership inference, and model inversion. It distinguishes between how explanations are obtained—through target releases, attacker-derived methods, secondary disclosure, privileged access, or global artifacts—and shows that explanations can lower extraction costs and reveal membership signals via statistics, recourse distance, and robustness. The authors compare threat models, signals, and defenses, concluding that no single explanation type is always unsafe and that protection must be tailored to the specific acquisition path and target asset.

By Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday
arXiv AI
Jul 17

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

arXiv:2607. 14315v1 Announce Type: cross Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score.

By Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, Jonh Soldatos
arXiv AI
Aug 24

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.

By Ye Chen, Weining Zhang
arXiv AI
Sep 23

When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating Systems

The paper discusses how large language model agents now act as privileged principals with kernel‑grade authority, yet lack the trusted mediation traditionally required for operating‑system security. It introduces a taxonomy that distinguishes between provenance‑based deterministic checks and content‑semantic checks, identifying a central mediation gap in distinguishing data from instruction and authorized from unauthorized actions. The authors argue that this gap creates an irreducible risk of undetected attacks whenever inputs and actions are not pre‑enumerated, and they propose defenses across runtime monitoring, architectural separation, and authorization while critiquing current evaluation practices. They extend the analysis to AI‑native operating systems where the model itself serves as the arbitration core, outlining design constraints, challenges, and a research agenda.

By Li Zhang, Yang Sun, Jie Shi
arXiv Machine Learning
Aug 31

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

The paper introduces a privacy‑preserving zk‑SNARK audit framework that uses adversarial‑style probes to detect logit drift between an approved large language model and a modified deployment. It offers three probe families—token‑based (black‑box), embedding‑based (gray‑box), and stress probes (partial white‑box)—allowing users to balance sensitivity, access, and cost. Experiments across LLM architectures and GPU platforms show token‑based probes achieve the highest mean sensitivity while remaining practical in a black‑box setting, with Groth16 proving times scaling modestly from 1.02 to 1.78 seconds and constant proof size.

By Cameron Wilding, Mina Shaker, Fatemeh Ganji