arXiv Computation and Language

AI Alignment through a Game-theoretic Lens: A Survey

The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.

arXiv AI
Jul 17

Align AI to Dynamic Human-AI Workflows

arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.

By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv AI
Sep 3

TUX: Measuring Human--AI Tacit Understanding

The paper introduces TUX, a Tacit Understanding Index that measures how similarly humans and large language models (LLMs) place concepts along subjective spectra in a task inspired by the game Wavelength. Using 241 human participants and 200 profile-conditioned LLM agents across four models, the study finds that human–agent pairs with similar traits achieve higher TUX scores, indicating that tacit alignment is linked to person-level characteristics. Regression analyses show that richer predictor sets—including individual traits, decision-making styles, and confidence—improve the explainability of TUX beyond simple trait-distance baselines.

By Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha
arXiv AI
Aug 26

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

The paper reinterprets AI alignment as a social choice problem, framing it as linear optimization over a convex impact space. This approach links alignment protocols to welfare outcomes, enabling the use of welfare economics and mechanism design tools. The authors demonstrate strategyproof, unanimous mechanisms like voting-by-issues and random-dictatorship, and derive alignment protocols that maximize utilitarian welfare while respecting harm constraints, validated on real human preference data.

By Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia