arXiv AI

Open Problems in Constitutional Preference Reconstruction

arXiv:2606. 30116v1 Announce Type: new Abstract: Pairwise preference data is widely used for training and evaluating language models (e.

arXiv AI
Aug 26

Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling

The paper proposes a transparent, user‑configurable rule for selecting arguments in deliberative polls, replacing opaque learned rankers. It formalises argument selection over bipolar justification sets, introduces seven civic recommender criteria, and presents a one‑hop reversed endorsement flow rule that meets them. Experiments on 17,000 simulated runs show the rule performs comparably to random on coverage but outperforms other methods on endorsement mass and robustness under adversarial pressure.

By Muntaser Syed, Markus Zanker, Marius Silaghi
Hugging Face Trending Papers
Jul 9

Validity of LLMs as data annotators: AMALIA on authority

A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size.

arXiv AI
Sep 2

The Constitutional Coverage Trilemma in AI Governance

The paper investigates how well the implicit value rankings encoded by frontier AI systems—termed constitutional institutions—meet human demand. By auditing 23 large language model archetypes and surveying 1,649 U.S. participants, the authors find that user demand spans all five values (safety, helpfulness, honesty, autonomy, equity) but the supply is narrow, covering only about 2% of the demand space, with no model prioritizing helpfulness or autonomy. They propose a sparse two‑vertex menu that substantially reduces regret compared to the full set of models and formalize these observations as a budgeted‑pluralism trilemma. whyItMatters":"The study reveals a significant mismatch between the values users prioritize and the values encoded by current AI models, highlighting the need for more diverse and aligned constitutional designs."

By Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse
arXiv AI
Sep 7

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.

By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
arXiv AI
Jul 8

Constitutional Governance in Metric Spaces

arXiv:2605. 13362v3 Announce Type: replace-cross Abstract: Computational social choice and algorithmic decision theory offer rich aggregation theory but no end-to-end process for egalitarian self-governance: aggregation, deliberation, amendment, and consensus are each considered in isolation, with key metric-space aggregators being NP-hard.

By Ehud Shapiro, Nimrod Talmon
arXiv Computation and Language
Sep 24

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

The paper introduces a sentence‑level benchmark for judging large language models’ ability to classify interpretive canons used by the German Federal Constitutional Court, based on Larenz’s framework. It operationalizes these canons as classification criteria, provides a dataset of court decisions annotated at the sentence level, and evaluates four LLMs with both expert hand‑written prompts and prompts optimized via Genetic‑Pareto. The results show mean F1 scores between 70.4 and 79.2, with grammatical interpretation being the easiest and systematic interpretation the hardest, and indicate that expert prompts already offer a strong baseline.

By Felix Ringe