arXiv:2609.36740v1 Announce Type: new
Abstract: Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Lear...
By Ren Kishimoto, Koichi Tanaka, Haruka Kiyohara, Yusuke Narita, Yasuo Yamamoto, Nobuyuki Shimizu, Yuta Saito
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
By Kosuke Iguchi, Ren Kishimoto
arXiv:2607. 20655v1 Announce Type: cross Abstract: Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production.
By Chenyu Zhang
arXiv:2605.11151v3 Announce Type: replace
Abstract: Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key ch...
By Andrew Choi, Wei Xu
arXiv:2601.13885v2 Announce Type: replace-cross
Abstract: Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluatio...
By Esma Balk{\i}r, Alice Pernthaller, Marco Basaldella, Jos\'e Hern\'andez-Orallo, Nigel Collier
Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e. g.
The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.
By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.
By Abhishek Divekar
arXiv:2606. 08679v1 Announce Type: cross Abstract: Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts.
By Bitya Neuhof, Yuval Benjamini
The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.
By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv:2607. 01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level.
By Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen
arXiv:2609.38860v1 Announce Type: cross
Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
By Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang