We’ve written a paper arguing that long-term AI safety research needs social scientists to ensure AI alignment algorithms succeed when actual humans are involved. Properly aligning advanced AI systems with human values requires resolving many uncertainties related to the psychology of human rationality, emotion, and biases.
The article "Math for AI safety: an invitation for mathematicians" calls for new mathematical tools to ensure AI remains understandable, controllable, and cooperative. It outlines specific mathematical fields—logic, game theory, probability, algebra, representation theory, analysis, and geometry—each paired with an open problem tailored for mathematicians without AI safety background. The piece invites researchers to contribute to designing AI that is legible, steerable, and aligned with human values.
By Lionel Levine
One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve developed an algorithm which can infer what humans want by being told which of two proposed behaviors is better.
arXiv:2408. 02379v2 Announce Type: replace-cross Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act.
By Benjamin Fresz, Vincent Philipp G\"obels, Safa Omri, Danilo Brajovic, Andreas Aichele, Janika Kutz, Jens Neuh\"uttler, Marco F. Huber
Google DeepMind researches AI's harmful manipulation risks across areas like finance and health, leading to new safety measures.