arXiv:2506. 10677v3 Announce Type: replace-cross Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline.
By Otmane Sakhi, Alexandre Gilotte, David Rohde
arXiv:2606. 18750v1 Announce Type: cross Abstract: A/B testing has become the gold standard for data-driven decision-making in large-scale online experimentation, providing critical guidance for feature launch, pricing optimization, and user experience enhancement.
By Yu Zhang, Bokui Wan, Yongli Qin, Jinyong Ma, Yifan Guo
arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.
By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig
arXiv:2511. 00802v2 Announce Type: replace-cross Abstract: With data-driven development now widely adopted, online A/B testing is an established method for measuring the effects of new technologies.
By Jie JW Wu, Ayanda Patrick Herlihy, Ahmad Saleem Mirza, Ali Afoud, Fatemeh Fard
arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.
By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo
arXiv:2608. 07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong.
By Guilin Zhang, Kai Zhao
arXiv:2603. 10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure.
By Ruicheng Ao, Hongyu Chen, Siyang Gao, Hanwei Li, David Simchi-Levi
arXiv:2606. 20820v2 Announce Type: replace Abstract: Can we trust evaluation scores to capture an LLM's true real-world performance?
By Zhijian Zhou, Zesheng Ye, Zhaorun Chen, Bo Li, Feng Liu
arXiv:2607. 14604v1 Announce Type: new Abstract: Online controlled experiments are the gold standard for hypothesis testing in online platforms.
By Olivier Jeunen
arXiv:2607. 25356v1 Announce Type: cross Abstract: Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is foundational for data-centric AI pipelines, yet exhaustive scans over millions of rows are prohibitively slow for near-real-time monitoring.
By Laure Berti-Equille
arXiv:2608. 16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail.
By Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou
arXiv:2605. 04954v2 Announce Type: replace-cross Abstract: Per-instance algorithm selection (PIAS) takes advantage of complementarity between a set of algorithms by deciding which algorithm to run on a given instance.
By Koen van der Blom, Diederick Vermetten