Probabilistic Linear Explanations
arXiv:2609. 19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.
arXiv:2604. 16689v2 Announce Type: replace Abstract: Masking-based post-hoc explanation methods, such as KernelSHAP and LIME, estimate local feature importance by querying a black-box model under randomized perturbations.
arXiv:2609. 19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.
arXiv:2603. 14894v3 Announce Type: replace-cross Abstract: Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model explanations.
arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.
Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a fundamental constraint: language is a compressed and capacity-limited interface for conveying task information.
arXiv:2606. 30372v1 Announce Type: new Abstract: Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias.
arXiv:2607. 06407v1 Announce Type: new Abstract: The XAI community has studied a wide range of queries and scores for explaining predictions of ML models.
The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².
arXiv:2607. 17232v1 Announce Type: cross Abstract: Classical rate-distortion (RD) theory has long established the fundamental limits of lossy compression by quantifying the minimum number of bits required to represent a source under a prescribed distortion constraint.
arXiv:2606. 02385v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control.
arXiv:2606. 26620v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features.
arXiv:2606. 14758v1 Announce Type: cross Abstract: As Vision-Language Models are increasingly deployed in safety-critical applications, the trustworthiness of their explanations becomes crucial.