arXiv AI By Zahra Khotanlou, Hashir Ahmed, Chenghao Tan, Ahmed Abdelaal, Amir-Hossein Karimi

RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation

Read the original on arXiv AI →

arXiv:2606. 16113v1 Announce Type: new Abstract: Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

Explanations-Driven Active Feature Acquisition for Algorithmic Recourse

The paper introduces Explanation-Driven Feature Acquisition (EDFA), a method that jointly optimizes algorithmic recourse and feature acquisition by selecting features based on explanatory value per unit cost. Using Markov Blanket theory, EDFA unifies various explanation types and provides distribution‑free validity guarantees for recourse derived from partial information. Experiments on seven datasets show that EDFA requires fewer features than existing active feature acquisition baselines while maintaining accuracy and producing more actionable recourse.

By Vinura Galwaduge, Jagath Samarabandu
arXiv AI
4d ago

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.

By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang
arXiv AI
Aug 24

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.

By Ye Chen, Weining Zhang