arXiv Machine Learning

Certification of Machine Learning Models via Directional Sharpness

arXiv:2606. 25004v1 Announce Type: new Abstract: In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality.

arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody
arXiv AI
Aug 24

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

The paper argues that prediction‑based certifications—such as accuracy, calibration, and conformal coverage—are insufficient to guarantee trustworthy AI. It proves a separation theorem showing that a model can appear reliable under all prediction‑side certificates yet differ arbitrarily in explanation fidelity and deployment behaviour. The authors propose a competence envelope framework that combines both prediction and explanation certification to detect such hidden failures.

By Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis
arXiv Machine Learning
Jul 27

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.

By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
arXiv Machine Learning
Sep 23

Optimizing Canaries for Privacy Auditing with Metagradient Descent

The paper investigates black-box privacy auditing for differentially private learning algorithms, focusing on DP‑SGD. It introduces a method that optimizes the auditor’s canary set using metagradient descent, improving empirical lower bounds on privacy parameters compared to prior canary designs. The approach is shown to be DP‑SGD agnostic and efficient, with optimized canaries for small models remaining effective for larger DP‑SGD models.

By Matteo Boglioni, Terrance Liu, Andrew Ilyas, Zhiwei Steven Wu