arXiv:2508. 09219v3 Announce Type: replace-cross Abstract: Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies.
By Wilder Baldwin, Sepideh Ghanavati, Manuel Woersdoerfer
arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.
By Guillermo Del Pinal, Youngchan Lee, Min Ohn
The paper examines how users handle harmful value conflicts with AI companions. By analyzing 146 posts and conducting a week-long study with 22 participants using the Minion technology probe, the authors find that users blend softer and harder strategies, especially when conflicts involve Universalism and Tradition values. The study highlights that such conflicts create asymmetric responsibility, with users shouldering unilateral repair work that AI companions cannot reciprocate, suggesting a need for platform-level safeguards.
By Qing Xiao, Xianzhe Fan, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen
The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.
By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
The paper argues that AI should be evaluated not only by principles but by concrete protocols that translate commitments into roles, requirements, records, oversight, and assessment. It introduces a rupture test linking institutional baselines to system evaluation, and distinguishes evidence‑bounded deployment from measurement‑bounded governance. The authors propose the RISE AI architecture to make bounded, evidence‑based claims about Responsibility, Inclusivity, Safety, and Empowerment, emphasizing the need for engineering, institutional repair, and ongoing moral judgment.
By Nitesh V. Chawla, Paulo Benanti
The paper examines how the rise of AI capable of moral reasoning could reshape meta-ethics, traditionally focused on human ethics. It proposes a framework that identifies new questions about AI’s own ethics from both human and AI perspectives, dividing them into four domains. The author explores how existing meta-ethical theories might apply to these domains and argues that many human-centered formulations will need significant revision to accommodate AI.
By Shang Lu