Nonparametric LLM Evaluation from Preference Data
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
arXiv:2604. 17805v2 Announce Type: replace-cross Abstract: Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pairwise comparisons.
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
arXiv:2508. 11847v4 Announce Type: replace-cross Abstract: We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data.
arXiv:2608. 08422v1 Announce Type: cross Abstract: Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback.
arXiv:2606. 27997v1 Announce Type: new Abstract: Benchmarks of machine learning models often include many datasets, making evaluation expensive.
arXiv:2510. 06732v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as rerankers in information retrieval, yet their ranking behavior can be steered by small, natural-sounding prompts.
arXiv:2605. 29107v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity.
arXiv:2505. 23437v2 Announce Type: replace-cross Abstract: Ranking systems influence decision-making in high-stakes domains like health, education, and employment, where they can have substantial economic and social impacts.
arXiv:2601. 21817v2 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm.
arXiv:2606. 09623v1 Announce Type: new Abstract: When running marketing campaigns, retailers must decide which products to promote and which users to target.
arXiv:2608. 10045v1 Announce Type: cross Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models.
arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.
arXiv:2607. 01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level.