How to Correctly Report LLM-as-a-Judge Evaluations
arXiv:2511. 21140v4 Announce Type: replace Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators.
arXiv:2511. 21140v4 Announce Type: replace Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators.
arXiv:2508. 13187v4 Announce Type: replace-cross Abstract: Homelessness is a persistent social challenge, impacting millions worldwide.
arXiv:2609.13148v1 Announce Type: cross Abstract: Large language models are increasingly deployed as synthetic consumer panels, promising $97\%$ cost reductions over traditional surveys. Yet aggregat...
arXiv:2604. 01925v2 Announce Type: replace-cross Abstract: Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly.
arXiv:2605. 11954v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical designs.
arXiv:2607. 14888v1 Announce Type: cross Abstract: Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains.
The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.
arXiv:2602. 18518v2 Announce Type: replace Abstract: Content safety teams need metrics that reflect what users actually experience, not only what is reported.
The paper argues that computational text‑based ideal point estimation (CT‑IPE) methods should be viewed as configurable measurement pipelines rather than fixed estimators. It presents a large‑scale comparative experiment involving 17 CT‑IPE algorithms, 5,537 runs, and about 4.25 million left‑right position estimates, and describes shared infrastructure that enables joint execution of these heterogeneous methods. Sensitivity analyses reveal that most algorithms exhibit low hyperparameter sensitivity (ICC < .10), with any remaining sensitivity concentrated in a few key researcher choices such as the language or embedding model, seed keyword lists, and number of topics.
arXiv:2609.00222v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how...
arXiv:2501. 02211v2 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias is stable or an artifact of inference settings has only been studied in single proprietary models.
arXiv:2606. 00369v1 Announce Type: cross Abstract: Safe global deployment of AI models requires alignment with human values that vary across cultures.