The paper proposes a two‑stage framework, SL+LHF, that first learns low‑dimensional representations from noisy labeled data and then refines model alignment using human comparison feedback via a probabilistic bisection approach. It introduces the label‑noise‑to‑comparison‑accuracy (LNCA) ratio to theoretically identify when this framework outperforms pure supervised learning, showing that trading labels for comparisons reduces sample complexity when labels are scarce. Experiments on a high‑dimensional crowdfunding prediction task and an Amazon Mechanical Turk study confirm that incorporating human or large language model evaluators improves accuracy under a fixed query budget.
By Junyu Cao, Mohsen Bayati
arXiv:2602. 01658v2 Announce Type: replace-cross Abstract: Bandit algorithms have recently emerged as a powerful tool for evaluating machine learning models, including generative image models and large language models, by efficiently identifying top-performing candidates without exhaustive comparisons.
By Seyed Mohammad Hadi Hosseini, Amir Najafi, Mahdieh Soleymani Baghshah
arXiv:2510. 19119v2 Announce Type: replace Abstract: In networked environments, it is common for users to share recommendations about content, products, services, and possible courses of action.
By Ahmed Sayeed Faruk, Mohammad Shahverdikondori, Elena Zheleva
arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
By Yan Dai, Negin Golrezaei, Patrick Jaillet
LLMAR is a tuning‑free recommendation framework designed for sparse, text‑rich industrial B2B domains. It transforms user behavioral history into structured semantic motives using LLM inference, employs a reflection loop to self‑correct hallucinations, and operates cost‑effectively with asynchronous batch processing. Experiments on MovieLens‑1M, Amazon Prime Pantry, and a construction risk dataset show LLMAR surpasses state‑of‑the‑art learning models, achieving up to a 54.6% nDCG@10 improvement while keeping inference costs around $1 per 1,000 users.
By Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda
arXiv:2004. 06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.
By Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou