arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.
By Guillermo Del Pinal, Youngchan Lee, Min Ohn
arXiv:2605. 08426v2 Announce Type: replace-cross Abstract: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety.
By Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Sch\"olkopf, Emanuele La Malfa, Zhijing Jin
arXiv:2511. 04177v2 Announce Type: replace Abstract: Personal AI agents are increasingly deployed in shared environments, where their actions affect not just the primary user they are assisting, but bystanders who never consented to being affected by the system.
By Claire Yang, Claire Jie Zhang, Maya Cakmak, Max Kleiman-Weiner
arXiv:2607. 00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized.
By Max Kanwal, Caryn Tran
arXiv:2607. 26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability.
By Sivasathivel Kandasamy
arXiv:2606. 08832v1 Announce Type: new Abstract: Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity.
By David Thorstad
arXiv:2608. 11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act.
By Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca
arXiv:2607. 18239v1 Announce Type: new Abstract: Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk.
By Mana Azarm, Qiyao Wei, Rahul Nambiar
arXiv:2607. 09766v1 Announce Type: new Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards.
By Yaowen Ye, Jacob Steinhardt
arXiv:2607. 12755v1 Announce Type: cross Abstract: AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities.
By Nathan G. Wood, Andrew P. Rebera
AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities. I.
arXiv:2606. 12420v1 Announce Type: cross Abstract: Our concepts of survival and self-interest were built for single, continuous biological lives.
By Dan Hendrycks