arXiv:2609.39025v1 Announce Type: cross
Abstract: Fairness assessment in algorithmic decisions that affect individuals, such as credit scoring, often relies on parity measures calculated at the aggre...
By Dalia Atif, Paolo Giudici
arXiv:2606. 26369v1 Announce Type: cross Abstract: Scoring functions are used to represent the relevance of individual documents.
By Shubham Singh, Ian A. Kash, Mesrob I. Ohannessian
arXiv:2608. 09899v1 Announce Type: new Abstract: In fair ranked link prediction, demographic parity ($\Delta_\mathrm{DP}$) is a common fairness metric.
By Valentijn Oldenburg, Floris de Kam, Stef de Wildt, Jarno Nilson Balk
arXiv:2602.08589v2 Announce Type: replace
Abstract: PageRank (PR) is a fundamental algorithm in graph machine learning tasks. Owing to the increasing importance of algorithmic fairness, we consider t...
By Emmanouil Kariotakis, Aritra Konar
arXiv:2603. 04689v3 Announce Type: replace-cross Abstract: Fair top-$k$ selection, which ensures appropriate proportional representation of members from minority or historically disadvantaged groups among the top-$k$ selected candidates, has drawn significant attention.
By Guangya Cai
arXiv:2608. 05958v1 Announce Type: new Abstract: The paper addresses several ranking-dependent decision support methods.
By Vitaliy Tsyganok, Sergii Kadenko, Oleh Andriichuk
arXiv:2606. 07988v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on reward models to align their outputs with diverse user preferences.
By Xiaoyan Zhao, Haoting Ni, Yang Zhang, Chunyuan Zheng, Haoxuan Li, Fuli Feng
The paper argues that traditional probabilistic fairness metrics can miss significant disparities in the actual consequences of decisions. By introducing a utility-based framework, the authors show that a process can satisfy ε-fairness yet still be maximally unfair when utilities are considered. They apply this framework to college admissions and credit‑risk assessment, demonstrating that equalizing probabilities alone may mask unequal utility outcomes across groups.
By Tolulope Fadina, Thorsten Schmidt
arXiv:2601. 23221v2 Announce Type: replace Abstract: As acquiring reliable ground-truth labels is usually costly, or infeasible, crowdsourcing and aggregation of noisy human annotations is the typical resort.
By Gabriel Singer, Samuel Gruffaz, Olivier Vo Van, Nicolas Vayatis, Argyris Kalogeratos
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
By Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan Feuerriegel
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
By Kosuke Iguchi, Ren Kishimoto
arXiv:2407. 14766v4 Announce Type: replace-cross Abstract: This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability and transparency of corrective methods, and on the opposition between two fairness criteria, namely Demographic Parity and Equalized Odds.
By Thomas Souverain, Paul \'Egr\'e