On Avoiding Power-Seeking by Artificial Intelligence

06/23/2022
by   Alexander Matt Turner, et al.
0

We do not know how to align a very intelligent AI agent's behavior with human interests. I investigate whether – absent a full solution to this AI alignment problem – we can build smart AI agents which have limited impact on the world, and which do not autonomously seek power. In this thesis, I introduce the attainable utility preservation (AUP) method. I demonstrate that AUP produces conservative, option-preserving behavior within toy gridworlds and within complex environments based off of Conway's Game of Life. I formalize the problem of side effect avoidance, which provides a way to quantify the side effects an agent had on the world. I also give a formal definition of power-seeking in the context of AI agents and show that optimal policies tend to seek power. In particular, most reward functions have optimal policies which avoid deactivation. This is a problem if we want to deactivate or correct an intelligent agent after we have deployed it. My theorems suggest that since most agent goals conflict with ours, the agent would very probably resist correction. I extend these theorems to show that power-seeking incentives occur not just for optimal decision-makers, but under a wide range of decision-making procedures.

READ FULL TEXT
research
12/03/2019

Optimal Farsighted Agents Tend to Seek Power

Some researchers have speculated that capable reinforcement learning (RL...
research
04/13/2023

Power-seeking can be probable and predictive for trained agents

Power-seeking behavior is a key source of risk from advanced AI, but our...
research
06/16/2022

Is Power-Seeking AI an Existential Risk?

This report examines what I see as the core argument for concern about e...
research
06/27/2022

Parametrically Retargetable Decision-Makers Tend To Seek Power

If capable AI agents are generally incentivized to seek power in service...
research
11/05/2014

Ethical Artificial Intelligence

This book-length article combines several peer reviewed papers and new m...
research
09/01/2021

Impossibility Results in AI: A Survey

An impossibility theorem demonstrates that a particular problem or set o...
research
07/22/2019

A Sufficient Statistic for Influence in Structured Multiagent Environments

Making decisions in complex environments is a key challenge in artificia...

Please sign up or login with your details

Forgot password? Click here to reset