arXiv AI

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

arXiv:2607. 00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized.

arXiv AI
Jul 17

Align AI to Dynamic Human-AI Workflows

arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.

By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv Computation and Language
Aug 31

AI Alignment through a Game-theoretic Lens: A Survey

The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.

By Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
arXiv AI
Aug 26

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

The paper reinterprets AI alignment as a social choice problem, framing it as linear optimization over a convex impact space. This approach links alignment protocols to welfare outcomes, enabling the use of welfare economics and mechanism design tools. The authors demonstrate strategyproof, unanimous mechanisms like voting-by-issues and random-dictatorship, and derive alignment protocols that maximize utilitarian welfare while respecting harm constraints, validated on real human preference data.

By Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
arXiv AI
Sep 12

Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

The paper proposes a developmental framework for autonomous artificial agents that emphasizes learning social norms and alignment through direct interaction with dynamic environments. It argues that intrinsic motivations such as curiosity and competence can guide exploration, but also complicate alignment with human goals. By drawing parallels to child development, the authors suggest that regulatory sandboxes serve as pedagogical spaces where agents gradually acquire moral agency and adapt their behaviors through experience and cooperation.

By Marica Notte, Ludovica Marinucci, Vieri Giuliano Santucci