arXiv Machine Learning

Is Model Instability just Noise to be Tolerated or a Property that can be Managed?

arXiv:2607. 10420v1 Announce Type: cross Abstract: In software analytics, rerunning the same analysis twice often yields different models and conclusions.

arXiv AI
6d ago

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

The paper introduces a unified evaluation protocol for robust counterfactual explanations (CFE), testing six robust methods and two baselines across four tabular datasets under eight types of model change. It shows that robustness scores vary by change type and that methods designed for one change family may not transfer to others, with RobX performing most consistently. The study emphasizes the need for a common protocol that defines model changes, measures their impact, and separates generation performance from robustness.

By Marcin Kostrzewa, Maciej Zi\k{e}ba
arXiv Machine Learning
Sep 1

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.

By Ahmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou, Shaina Raza