More Efficient Policy Learning via Optimal Retargeting

06/20/2019
by   Nathan Kallus, et al.
0

Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different actions, which can lead to unwieldy policy evaluation and poorly performing learned policies. We study a solution to this problem based on retargeting, that is, changing the population on which policies are optimized. We first argue that at the population level, retargeting may induce little to no bias. We then characterize the optimal reference policy centering and retargeting weights in both binary-action and multi-action settings. We do this in terms of the asymptotic efficient estimation variance of the new learning objective. We further consider bias regularization. Extensive empirical results in a simulation study and a case study of targeted job counseling demonstrate that retargeting is a fairly easy way to significantly improve any policy learning procedure.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/06/2020

Efficient Evaluation of Natural Stochastic Policies in Offline Reinforcement Learning

We study the efficient off-policy evaluation of natural stochastic polic...
research
12/02/2021

Generalizing Off-Policy Learning under Sample Selection Bias

Learning personalized decision policies that generalize to the target po...
research
05/02/2023

Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare

Many reinforcement learning (RL) applications have combinatorial action ...
research
11/22/2021

Case-based off-policy policy evaluation using prototype learning

Importance sampling (IS) is often used to perform off-policy policy eval...
research
10/09/2020

Discussion of Kallus (2020) and Mo, Qi, and Liu (2020): New Objectives for Policy Learning

We discuss the thought-provoking new objective functions for policy lear...
research
03/04/2022

Interpretable Off-Policy Learning via Hyperbox Search

Personalized treatment decisions have become an integral part of modern ...
research
06/06/2023

Fair and Robust Estimation of Heterogeneous Treatment Effects for Policy Learning

We propose a simple and general framework for nonparametric estimation o...

Please sign up or login with your details

Forgot password? Click here to reset