arXiv AI

Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis

arXiv:2508. 00129v2 Announce Type: replace Abstract: Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-documented threat to the reliability of Multi-Criteria Decision Analysis (MCDA) methods.

arXiv Machine Learning
Aug 4

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.

By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao
arXiv AI
Sep 25

Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes

Augur is a synthetic decision laboratory that simulates how users will react to product and policy changes before they are released. It constructs a typed knowledge graph from change documents, populates a persona market, runs simulations, and produces an auditable decision memo recommending one of five actions. Using a dataset of 50 real episodes (Gold‑50), the authors evaluate the system’s five‑way release verdicts and find that evaluation design, rather than model capability, largely drives performance differences among frontier and open‑weight models.

By Rahul Khedar, Mayank Malhotra, Avinash Karn
arXiv Machine Learning
Sep 22

Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints

The paper introduces a counterfactual tool ranking framework that accounts for authority, historical support, and estimation nuances. Using eleven enterprise-inspired tools, synthetic and real-world experiments on the Berkeley Function Calling Leaderboard, the study compares direct regression and doubly robust (DR) methods, finding that DR performs better in shifted environments while direct regression excels in linear settings. The authors also evaluate Qwen2.5 models on held-out tasks, analyze policy differences under missing support, and present a falsifiable evaluation method with publicly available evidence.

By Jiapeng Li