arXiv AI By Darrell Lewis-Sandy

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

Read the original on arXiv AI →

arXiv:2606. 28710v1 Announce Type: new Abstract: We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

The Constitutional Coverage Trilemma in AI Governance

The paper investigates how well the implicit value rankings encoded by frontier AI systems—termed constitutional institutions—meet human demand. By auditing 23 large language model archetypes and surveying 1,649 U.S. participants, the authors find that user demand spans all five values (safety, helpfulness, honesty, autonomy, equity) but the supply is narrow, covering only about 2% of the demand space, with no model prioritizing helpfulness or autonomy. They propose a sparse two‑vertex menu that substantially reduces regret compared to the full set of models and formalize these observations as a budgeted‑pluralism trilemma. whyItMatters":"The study reveals a significant mismatch between the values users prioritize and the values encoded by current AI models, highlighting the need for more diverse and aligned constitutional designs."

By Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse
arXiv AI
Sep 15

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

The paper studies how to safely delegate action approval to multiple AI reviewers when the reviewers themselves may be misaligned. It introduces a weaker condition—k‑robust coalitional alignment—under which a threshold rule that tolerates up to k disapprovals guarantees that the principal’s expected utility is at least as good as a baseline policy. The authors extend this characterization to sequential decision‑making in discounted MDPs and show that full‑panel coverage of reward functions ensures safety in Nash equilibria, while more permissive thresholds can lead to unsafe outcomes. Experiments demonstrate that collective review can remain sound even when individual reviewers are not fully aligned, provided some disapprovals are allowed.

By Natalie Collina, Surbhi Goel, Aaron Roth, Sikata Bela Sengupta
arXiv AI
Sep 3

The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

The paper examines how an agent’s probability report is evaluated twice—once by a strictly proper scoring rule and again by an approval rule that determines a decision. It shows that when the approval rule is welfare‑maximizing, it cannot be affine, yet the resulting distortion is predictable and can be mitigated by a reserve report that neutralizes the cost of pretending to be the marginal type. A Lipschitz rule with a single kink achieves first‑best welfare, while smooth rules cannot, and the key constraint is the steepness of the rule rather than its smoothness.

By Lauri Lov\'en, Sasu Tarkoma
arXiv AI
Jun 2

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

arXiv:2606. 00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts.

By Rana Muhammad Usman