arXiv AI By David Thorstad

Instrumental convergence and power-seeking

Read the original on arXiv AI →

arXiv:2606. 08832v1 Announce Type: new Abstract: Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 15

A Virtuous AI is an Existential Risk

arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.

By Guillermo Del Pinal, Youngchan Lee, Min Ohn
OpenAI Blog
Feb 20, 2018

Preparing for malicious uses of AI

We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.

arXiv AI
Sep 10

Agentic Inequality

arXiv:2510.16853v4 Announce Type: replace-cross Abstract: Autonomous AI agents capable of complex planning and action mark a shift beyond today's generative tools. As these systems enter political an...

By Matthew Sharp, Omer Bilgin, Iason Gabriel, Lewis Hammond