The paper introduces Explanation-Driven Feature Acquisition (EDFA), a method that jointly optimizes algorithmic recourse and feature acquisition by selecting features based on explanatory value per unit cost. Using Markov Blanket theory, EDFA unifies various explanation types and provides distribution‑free validity guarantees for recourse derived from partial information. Experiments on seven datasets show that EDFA requires fewer features than existing active feature acquisition baselines while maintaining accuracy and producing more actionable recourse.
By Vinura Galwaduge, Jagath Samarabandu
arXiv:2607. 22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome.
By Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.
By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang
arXiv:2606. 29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof.
By Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman
The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.
By Ye Chen, Weining Zhang
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu