arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.
By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.
By Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
arXiv:2609.01595v1 Announce Type: cross
Abstract: We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unkn...
By Dirk Bergemann, Andrew Koh, Stephen Morris
arXiv:2604. 14990v2 Announce Type: replace Abstract: The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem.
By Till Mossakowski, Helena Esther Grass
arXiv:2608. 10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values?
By Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
The paper reinterprets AI alignment as a social choice problem, framing it as linear optimization over a convex impact space. This approach links alignment protocols to welfare outcomes, enabling the use of welfare economics and mechanism design tools. The authors demonstrate strategyproof, unanimous mechanisms like voting-by-issues and random-dictatorship, and derive alignment protocols that maximize utilitarian welfare while respecting harm constraints, validated on real human preference data.
By Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia