Probabilistic model checking for Markov decision processes (MDPs) provides quantitative guarantees, but often offers limited insight into why undesired outcomes occur. Probability-raising (PR) causality addresses this by identifying states whose visitation increases the probability of reaching designated states.
arXiv:2607. 26787v1 Announce Type: new Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations.
By Jule Schmidt, Maximilian Weininger, Clemens Dubslaff, David Parker, Nils Jansen
arXiv:2511. 19849v2 Announce Type: replace-cross Abstract: Recurrence objectives, where a target region must be visited infinitely often, are a fundamental class of specifications for Markov decision processes (MDPs) and form the core of $\omega$-regular and linear temporal logic (LTL) objectives.
By Dominik Wagner, Leon Witzman, Luke Ong
arXiv:2604.01024v2 Announce Type: replace
Abstract: We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). A common approach to le...
By Philip Jordan, Maryam Kamgarpour
The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.
By Lu-Chin Chang, Suguman Bansal
arXiv:2606. 00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs).
By Edwin Hamel-De le Court, Thom Badings, Alessandro Abate, Francesco Belardinelli, Francesco Fabiano
arXiv:2602. 23545v2 Announce Type: replace Abstract: In the real world, planning is often challenged by distribution shifts.
By Matteo Ceriscioli, Karthika Mohan
arXiv:2609.07461v1 Announce Type: new
Abstract: We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. T...
By Jan Corazza, Daniil Kaminskyi, Simon Lutz, Patrick Nossol, Hadi Partovi Aria, Zhe Xu, Daniel Neider
The paper introduces Probabilistic Causal Impact (PCI), a framework that blends actual causality (AC) with Pearl’s probability of necessity and sufficiency to provide tractable, causally grounded explanations. PCI reframes explainability as an estimation problem on a probabilistic causal model, enabling efficient approximation via Monte Carlo sampling. The authors evaluate PCI on synthetic and real-world data, demonstrating consistency with AC, scalability, and applicability to complex continuous systems and large-scale causal machine learning models.
By Rafal Urbaniak, Sam Witty, Daniel Waxman, Andy Zane, Poorva Garg, Emily Bunnapradist, Sankaran Vaidyanathan, Jack Feser, Drew Lehe, Eli Bingham
arXiv:2505. 15274v4 Announce Type: replace Abstract: Probabilities of causation (PoCs) are fundamental quantities for counterfactual analysis and personalized decision making.
By Xin Shu, Shuai Wang, Ang Li
arXiv:2603. 06946v2 Announce Type: replace Abstract: Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority.
By Ege C. Kaya, Mahsa Ghasemi, Abolfazl Hashemi
arXiv:2608. 07230v1 Announce Type: new Abstract: Probabilistic logic programming is a formalism of statistical relational artificial intelligence that supports causal queries, including interventions from outside the system.
By Zora Wurm, Kilian R\"uckschlo{\ss}, Felix Weitk\"amper