arXiv Machine Learning By Koki Konishi, Masataka Ushiku, Yuta Saito

A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation Methods

Read the original on arXiv Machine Learning →

arXiv:2607. 01958v1 Announce Type: new Abstract: A/B testing is the gold standard for selecting the better algorithm in online services.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 3

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.

By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig
arXiv Machine Learning
Jul 9

Best-Arm Identification with Generative Proxy

arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.

By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo