arXiv Machine Learning By Binshuang Li

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

Read the original on arXiv Machine Learning →

arXiv:2608. 12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints

The paper introduces a counterfactual tool ranking framework that accounts for authority, historical support, and estimation nuances. Using eleven enterprise-inspired tools, synthetic and real-world experiments on the Berkeley Function Calling Leaderboard, the study compares direct regression and doubly robust (DR) methods, finding that DR performs better in shifted environments while direct regression excels in linear settings. The authors also evaluate Qwen2.5 models on held-out tasks, analyze policy differences under missing support, and present a falsifiable evaluation method with publicly available evidence.

By Jiapeng Li