Nonparametric LLM Evaluation from Preference Data
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
arXiv:2608. 08422v1 Announce Type: cross Abstract: Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback.
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
arXiv:2602. 07298v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are crucial for guiding research and optimizing resource allocation.
arXiv:2310. 11714v5 Announce Type: replace Abstract: Ranking generative models based on the fidelity and diversity of their outputs is required to identify the best generator in a group of candidate generative AI models.
arXiv:2606. 04603v1 Announce Type: cross Abstract: Approximate Nearest Neighbour search indices form the backbone of real-world recommender systems, enabling real-time candidate retrieval over million-item catalogues.
arXiv:2607. 20737v1 Announce Type: new Abstract: Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems.
arXiv:2511. 08307v2 Announce Type: replace-cross Abstract: Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries.
arXiv:2604. 17805v2 Announce Type: replace-cross Abstract: Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pairwise comparisons.
arXiv:2608. 03647v1 Announce Type: cross Abstract: Out-of-distribution (OOD) recommendation is vulnerable to preference shifts induced by a latent environment.
arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
arXiv:2606. 06779v1 Announce Type: cross Abstract: In multi-vertical e-commerce platforms like DoorDash, relatively newer product verticals such as grocery and retail present a significant opportunity for personalization innovation.
arXiv:2605. 17779v2 Announce Type: replace Abstract: Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs).
arXiv:2606. 07492v1 Announce Type: cross Abstract: The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale.