Probabilistic Linear Explanations
arXiv:2609. 19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.
arXiv:2603. 14894v3 Announce Type: replace-cross Abstract: Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model explanations.
arXiv:2609. 19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.
arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
Interpretable AI with Local Distillation proposes a method where a black‑box teacher model guides a regularized linear student model at each query point. The teacher defines locality by upweighting training observations with similar predicted outcomes and anchors the fit with its own prediction at the query point, treated as a pseudo‑observation. By adding Gaussian randomization and refitting, the approach identifies reliable features and stable subgroups, achieving near‑teacher accuracy while producing sparse, locally interpretable linear models.
The paper introduces XCal-FL, a federated learning algorithm that dynamically calibrates differential privacy noise using three signals—prediction logit variations, counterfactual margins, and saliency concentration—to improve both predictive accuracy and explanation fidelity. Experiments on medical imaging datasets demonstrate that XCal-FL outperforms static-noise and state‑of‑the‑art adaptive DP methods, achieving over 10% better accuracy and up to fivefold higher explanation fidelity while using privacy budgets more efficiently. The study highlights that explanation fidelity behaves non‑linearly with privacy loss, indicating that explainability is a separate dimension of the privacy trade‑off.
arXiv:2604. 16689v2 Announce Type: replace Abstract: Masking-based post-hoc explanation methods, such as KernelSHAP and LIME, estimate local feature importance by querying a black-box model under randomized perturbations.
arXiv:2605.30327v2 Announce Type: replace-cross Abstract: Frontier reasoning models are produced by post-training base language models with reinforcement learning. Recent work has challenged this by...
arXiv:2605. 27618v2 Announce Type: replace Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may not always be reliable.
The paper introduces XCal-FL, a federated learning algorithm that dynamically calibrates differential privacy noise using three signals—prediction logit variations, counterfactual margins, and saliency concentration—to improve both predictive accuracy and explanation fidelity. Experiments on medical imaging datasets demonstrate that XCal-FL outperforms static-noise and state‑of‑the‑art adaptive DP methods, achieving over 10% better accuracy and up to five‑fold higher explanation fidelity while using privacy budgets more efficiently. The study reveals that explanation fidelity behaves non‑linearly with privacy loss, indicating that explainability is a separate dimension of the privacy trade‑off that cannot be inferred from utility alone.
arXiv:2608. 25897v1 Announce Type: new Abstract: Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior.
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
arXiv:2604. 04535v2 Announce Type: replace Abstract: Modern machine learning systems, such as generative models and recommendation systems, often evolve through a cycle of deployment, user interaction, and periodic model updates.
Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods...