arXiv:2608. 08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way.
By Jobst Heitzig, Ram Potham
arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.
By Guillermo Del Pinal, Youngchan Lee, Min Ohn
arXiv:2608. 05173v1 Announce Type: cross Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole.
By Peter Barnett
We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.
arXiv:2608. 08022v1 Announce Type: new Abstract: Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant criminal or negligent damage.
By Mark Burgess
arXiv:2510.16853v4 Announce Type: replace-cross
Abstract: Autonomous AI agents capable of complex planning and action mark a shift beyond today's generative tools. As these systems enter political an...
By Matthew Sharp, Omer Bilgin, Iason Gabriel, Lewis Hammond