arXiv:2607. 10038v1 Announce Type: cross Abstract: For many years, the pairwise comparison method has been widely used for decision-making involving experts.
By Konrad Ku{\l}akowski, Jacek Szybowski
arXiv:2604.17805v3 Announce Type: replace-cross
Abstract: Maximum-likelihood pairwise ranking is a com- mon computational mechanism for prioritization, reputation estimation, and comparison-driven de...
By Junyi Yao, Zihao Zheng, Jiayu Long
arXiv:2606. 17756v1 Announce Type: new Abstract: Fairness has become a central concern in ranking problems involving individuals or social groups, particularly under the Responsible Artificial Intelligence agenda.
By Guilherme Dean Pelegrina, Renata Pelissari
arXiv:2609.16779v1 Announce Type: new
Abstract: LLMs are increasingly employed in a wide range of decision-making tasks. However, the opacity of their internal reasoning makes it difficult to validat...
By Han Zhiguang (IRIT-MELODI, UT3, IPAL), Farah Benamara (IRIT-MELODI, UT3, IPAL), Pascale Zarat\'e (IRIT, UT Capitole, IRIT-ADRIA)
arXiv:2609.07785v1 Announce Type: new
Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that c...
By Wei-Jung Huang
arXiv:2508. 00129v2 Announce Type: replace Abstract: Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-documented threat to the reliability of Multi-Criteria Decision Analysis (MCDA) methods.
By Juan Bautista Cabral, Gonzalo Giarda, Diego Nicol\'as Gimenez Irusta, Paula Pacheco, Alvaro Roy Schachner, Agust\'in Borda
arXiv:2607. 16259v1 Announce Type: new Abstract: Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks.
By Bitya Neuhof, Yuval Benjamini
arXiv:2606. 08679v1 Announce Type: cross Abstract: Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts.
By Bitya Neuhof, Yuval Benjamini
arXiv:2606. 07253v1 Announce Type: new Abstract: Traditional TOPSIS derives its reference points -- the Positive Ideal Solution ($PIS$) and Negative Ideal Solution ($NIS$) -- from the observed alternative set, making rankings susceptible to misalignment with decision-maker (DM) requirements, sensitivity to outlier performances, and rank reversal.
By Leonardo Fernandes Costa, Helder Gomes Costa, Diogo Lima, Brunno Rodrigues
arXiv:2606. 30412v1 Announce Type: cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources.
By Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das
arXiv:2606. 29159v1 Announce Type: new Abstract: Offline root-cause-analysis (RCA) benchmarks commonly rank methods by a single pooled top-1 accuracy across multiple subsystems, and engineers often read the pooled winner as a recommendation for their own subsystem.
By Lining Hu, Ting Liu, Yuzhuo Fu
Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.
By Bruno Brocai, Maria Becker