The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.
By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
The paper reports the first empirical study comparing how humans and large language models (LLMs) evaluate perceived moral agency (PMA) in both human and autonomous artificial agents within smart city scenarios. Using a validated PMA scale, 190 human participants and various LLMs were assessed, revealing that humans are perceived to have higher moral agency than artificial agents. When confronted with moral dilemmas, LLMs focus on situational factors such as harm severity and urgency, mirroring the context‑sensitivity observed in human raters.
By Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen
arXiv:2510. 16380v2 Announce Type: replace-cross Abstract: As AI systems progress, we rely more on them to make decisions with us and for us.
By Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Rapha\"el Milli\`ere, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
arXiv:2606. 00013v1 Announce Type: cross Abstract: Social conformity is a well-documented phenomenon in which individuals shift their opinions towards those of a social majority.
By Yana Venerina, Dmitry Koch, Nare Meloyan, Gerda Prutko, Valeriia Lelik, Victoria Taova, Andrey Kurpatov
arXiv:2606. 28345v1 Announce Type: cross Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first.
By Carmen Ng, Gjergji Kasneci
arXiv:2608. 14522v1 Announce Type: new Abstract: As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation.
By Taenyun Kim, Edyta Bogucka, Daniele Quercia
arXiv:2608. 10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values?
By Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
arXiv:2606. 11635v1 Announce Type: cross Abstract: For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly.
By Menghang Zhu, Seth Lazar
arXiv:2604. 14990v2 Announce Type: replace Abstract: The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem.
By Till Mossakowski, Helena Esther Grass
arXiv:2608. 08220v1 Announce Type: new Abstract: The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways.
By Aleks Knoks, Marija Slavkovik
arXiv:2608. 08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict.
By Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
arXiv:2606. 15507v1 Announce Type: new Abstract: Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it.
By Ali Dasdan, Manan Shah, W. Russell Neuman, Chad Coleman, Kund Meghani, Safinah Ali