Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

08/13/2019
by   Tom Everitt, et al.
3

Can an arbitrarily intelligent reinforcement learning agent be kept under control by a human user? Or do agents with sufficient intelligence inevitably find ways to shortcut their reward signal? This question impacts how far reinforcement learning can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence. In this paper, we use an intuitive yet precise graphical model called causal influence diagrams to formalize reward tampering problems. We also describe a number of tweaks to the reinforcement learning objective that prevent incentives for reward tampering. We verify the solutions using recently developed graphical criteria for inferring agent incentives from causal influence diagrams.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/26/2019

Understanding Agent Incentives using Causal Influence Diagrams, Part I: Single Action Settings

Agents are systems that optimize an objective function in an environment...
research
02/03/2022

Reward is not enough: can we liberate AI from the reinforcement learning paradigm?

I present arguments against the hypothesis put forward by Silver, Singh,...
research
05/13/2021

Intelligence and Unambitiousness Using Algorithmic Information Theory

Algorithmic Information Theory has inspired intractable constructions of...
research
10/19/2018

Intrinsic Social Motivation via Causal Influence in Multi-Agent RL

We derive a new intrinsic social motivation for multi-agent reinforcemen...
research
02/02/2021

Agent Incentives: A Causal Perspective

We present a framework for analysing agent incentives using causal influ...
research
02/27/2023

Reinforcement Learning with Depreciating Assets

A basic assumption of traditional reinforcement learning is that the val...
research
07/25/2019

Interactive Lungs Auscultation with Reinforcement Learning Agent

To perform a precise auscultation for the purposes of examination of res...

Please sign up or login with your details

Forgot password? Click here to reset