The paper examines how Preference Inference (PI) models used in large-scale participatory democracy platforms can alter the perceived consensus and minority support by predicting missing votes. It introduces a collective‑centric evaluation framework that assesses whether inferred votes maintain key properties of the overall preference landscape, rather than focusing solely on individual prediction accuracy. Using the largest multilingual dataset to date—four consultations with over 90,000 participants, 1 million votes, and 22 languages—the study finds that models with similar predictive accuracy can differ markedly in how well they preserve the collective structure, underscoring that accuracy alone is insufficient for evaluating PI in democratic contexts.
By Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner, Nazanin Shafiabadi, Laur\`ene Cave, David Mas, Jean-Philippe Cointet, Benjamin Piwowarski, Fran\c{c}ois Yvon
arXiv:2609.22133v1 Announce Type: new
Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at...
By Kentaro Nakamura, Jing Ling Tan, George Yean
arXiv:2401.08038v2 Announce Type: replace-cross
Abstract: A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and com...
By Wenjun Qiu, David Lie, Lisa Austin
Calpric is a system that combines automatic text selection, segmentation, active learning, and crowdsourced annotation to create a large, balanced training set for privacy policy classification. By simplifying the labeling task, it enables untrained crowd workers to match the performance of trained annotators and reduces inter‑annotator disagreement, cutting labeling costs. The approach yields a dataset of 16,000 policy text segments across nine data categories and produces models that deliver accurate, fine‑grained labels at a cost of roughly $0.92–$1.71 per segment.
By Wenjun Qiu, David Lie, Lisa Austin
arXiv:2603.15220v2 Announce Type: replace
Abstract: Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena. While prior studies have attempted...
By Minsung Cho, Jaehyung Kim
arXiv:2608.29943v1 Announce Type: new
Abstract: Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to re...
By Shicheng Hu, Runzhi Tian, Ziqiao Wang, Yongyi Mao