arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.
By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
By Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
arXiv:2607. 00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized.
By Max Kanwal, Caryn Tran
arXiv:2605. 01642v2 Announce Type: replace Abstract: Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time.
By Rachel Freedman
arXiv:2404. 02039v5 Announce Type: replace Abstract: Game environments provide rich, controllable settings that stimulate many aspects of real-world complexity.
By Sihao Hu, Tiansheng Huang, Gaowen Liu, Ramana Rao Kompella, Fatih Ilhan, Selim Furkan Tekin, Yichang Xu, Zachary Yahn, Ling Liu
arXiv:2608. 12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.
By Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg
arXiv:2608.18265v2 Announce Type: replace-cross
Abstract: We introduce a general, easy-to-implement AI-based method for modeling and analyzing the structure and complexity of human behavior. We assig...
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
The paper introduces TUX, a Tacit Understanding Index that measures how similarly humans and large language models (LLMs) place concepts along subjective spectra in a task inspired by the game Wavelength. Using 241 human participants and 200 profile-conditioned LLM agents across four models, the study finds that human–agent pairs with similar traits achieve higher TUX scores, indicating that tacit alignment is linked to person-level characteristics. Regression analyses show that richer predictor sets—including individual traits, decision-making styles, and confidence—improve the explainability of TUX beyond simple trait-distance baselines.
By Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha
arXiv:2608. 10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values?
By Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
arXiv:2512. 20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment.
By Anselm Paulus, Ilia Kulikov, Brandon Amos, R\'emi Munos, Ivan Evtimov, Kamalika Chaudhuri, Arman Zharmagambetov
The paper reinterprets AI alignment as a social choice problem, framing it as linear optimization over a convex impact space. This approach links alignment protocols to welfare outcomes, enabling the use of welfare economics and mechanism design tools. The authors demonstrate strategyproof, unanimous mechanisms like voting-by-issues and random-dictatorship, and derive alignment protocols that maximize utilitarian welfare while respecting harm constraints, validated on real human preference data.
By Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
arXiv:2510. 09330v3 Announce Type: replace Abstract: Ensuring that large language models (LLMs) comply with safety requirements is a central challenge in AI deployment.
By Tuan Nguyen, Long Tran-Thanh