The paper introduces Explanation-Driven Feature Acquisition (EDFA), a method that jointly optimizes algorithmic recourse and feature acquisition by selecting features based on explanatory value per unit cost. Using Markov Blanket theory, EDFA unifies various explanation types and provides distribution‑free validity guarantees for recourse derived from partial information. Experiments on seven datasets show that EDFA requires fewer features than existing active feature acquisition baselines while maintaining accuracy and producing more actionable recourse.
By Vinura Galwaduge, Jagath Samarabandu
arXiv:2607. 22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome.
By Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.
By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang
arXiv:2606. 29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof.
By Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman
The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.
By Ye Chen, Weining Zhang
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
arXiv:2608.30956v1 Announce Type: cross
Abstract: Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input...
By Mattia Cerrato, Otto Sahlgren, Xenia Heilmann
REFINE is a tool-agnostic, evidence-aware multi-agent approach that generates Java file-level refactoring candidates by combining static analysis, smell-informed planning, LLM-based transformation, and automated re-analysis. In experiments on 450 Java files from 15 open-source systems, REFINE reduced detected code smells by 68–73% across three LLM configurations, achieving higher median reductions with smaller edits compared to a direct-prompt baseline. However, the tool’s outputs still pose risks such as assert/fail-call changes and public-method removal, requiring compilation, testing, dependency analysis, and human review before deployment.
By Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson
The paper introduces a unified evaluation protocol for robust counterfactual explanations (CFE), testing six robust methods and two baselines across four tabular datasets under eight types of model change. It shows that robustness scores vary by change type and that methods designed for one change family may not transfer to others, with RobX performing most consistently. The study emphasizes the need for a common protocol that defines model changes, measures their impact, and separates generation performance from robustness.
By Marcin Kostrzewa, Maciej Zi\k{e}ba
RubricRefine is a training‑free pre‑execution refinement method that generates task‑specific rubrics from tool documentation, scores candidate code against explicit contract checks, and iteratively repairs failures before execution. It achieves an average score of 0.86 across seven models on M3ToolEval without any execution attempts, outperforming prior inference‑time baselines while incurring lower latency. The approach shows consistent performance on single‑step API‑Bank tasks and maintains an advantage in multi‑turn settings on AppWorld, with its effectiveness tied to the quality of the supplied documentation.
By Will LeVine, Brendan Evers, Sam Saltwick, Abhay Venkatesh
arXiv:2606. 18832v1 Announce Type: cross Abstract: Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems.
By K. Darshana Abeyrathna, Sara El Mekkaoui, Nils Enric Canut Taugb{\o}l, Anuja Vats
arXiv:2606. 08696v1 Announce Type: cross Abstract: Counterfactual recourse aims to provide actionable feature changes that would alter an unfavorable decision made by a predictive model.
By Yasuo Tabei