The paper investigates how well the implicit value rankings encoded by frontier AI systems—termed constitutional institutions—meet human demand. By auditing 23 large language model archetypes and surveying 1,649 U.S. participants, the authors find that user demand spans all five values (safety, helpfulness, honesty, autonomy, equity) but the supply is narrow, covering only about 2% of the demand space, with no model prioritizing helpfulness or autonomy. They propose a sparse two‑vertex menu that substantially reduces regret compared to the full set of models and formalize these observations as a budgeted‑pluralism trilemma.
whyItMatters":"The study reveals a significant mismatch between the values users prioritize and the values encoded by current AI models, highlighting the need for more diverse and aligned constitutional designs."
By Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse
The paper studies how to safely delegate action approval to multiple AI reviewers when the reviewers themselves may be misaligned. It introduces a weaker condition—k‑robust coalitional alignment—under which a threshold rule that tolerates up to k disapprovals guarantees that the principal’s expected utility is at least as good as a baseline policy. The authors extend this characterization to sequential decision‑making in discounted MDPs and show that full‑panel coverage of reward functions ensures safety in Nash equilibria, while more permissive thresholds can lead to unsafe outcomes. Experiments demonstrate that collective review can remain sound even when individual reviewers are not fully aligned, provided some disapprovals are allowed.
By Natalie Collina, Surbhi Goel, Aaron Roth, Sikata Bela Sengupta
arXiv:2610.00902v1 Announce Type: cross
Abstract: One way to make AI systems safe is to shape what the system is: its objective and dispositions. We take a complementary route: treat the agents' char...
By P. Jameson Graber
arXiv:2603. 13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget.
By Majid Ghasemi, Mark Crowley
The paper examines how an agent’s probability report is evaluated twice—once by a strictly proper scoring rule and again by an approval rule that determines a decision. It shows that when the approval rule is welfare‑maximizing, it cannot be affine, yet the resulting distortion is predictable and can be mitigated by a reserve report that neutralizes the cost of pretending to be the marginal type. A Lipschitz rule with a single kink achieves first‑best welfare, while smooth rules cannot, and the key constraint is the steepness of the rule rather than its smoothness.
By Lauri Lov\'en, Sasu Tarkoma
arXiv:2606. 00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts.
By Rana Muhammad Usman
The paper proposes a transparent, user‑configurable rule for selecting arguments in deliberative polls, replacing opaque learned rankers. It formalises argument selection over bipolar justification sets, introduces seven civic recommender criteria, and presents a one‑hop reversed endorsement flow rule that meets them. Experiments on 17,000 simulated runs show the rule performs comparably to random on coverage but outperforms other methods on endorsement mass and robustness under adversarial pressure.
By Muntaser Syed, Markus Zanker, Marius Silaghi
arXiv:2608.22444v1 Announce Type: new
Abstract: The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that rea...
By Isotta Magistrali, Chen Shani
arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.
By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
The paper introduces PRIMUS, an extension of the PRIMA framework that combines prime‑power agent identity with BLS aggregate signatures to improve governance in multi‑agent federations. PRIMUS achieves a safe‑kill threshold that eliminates false‑positive agent termination under noisy conditions, identifies an economic boundary where singleton governance outperforms Byzantine quorum, and implements VRF succession with lease and fencing for unconditional safety under partial synchrony. The study also explores converting PRIMA’s binary artifact‑fidelity verdict into a graded fitness signal, demonstrating strong calibration against injected faults but reduced effectiveness on real LLM‑generated candidates, and reports a program cost of USD 164.78.
By Sasank Annapureddy, Anjaneya Prasad Thamatani
arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan
arXiv:2606. 00235v1 Announce Type: cross Abstract: We argue that governance must transition from a normative discipline to an engineering discipline, and we develop a formal framework, inspired by the physics of metamaterials, to make this transition quantitative and testable.
By David Orban