Balanced Policy Evaluation and Learning

05/21/2017
by   Nathan Kallus, et al.
0

We present a new approach to the problems of evaluating and learning personalized decision policies from observational data of past contexts, decisions, and outcomes. Only the outcome of the enacted decision is available and the historical policy is unknown. These problems arise in personalized medicine using electronic health records and in internet advertising. Existing approaches use inverse propensity weighting (or, doubly robust versions) to make historical outcome (or, residual) data look like it were generated by a new policy being evaluated or learned. But this relies on a plug-in approach that rejects data points with a decision that disagrees with the new policy, leading to high variance estimates and ineffective learning. We propose a new, balance-based approach that too makes the data look like the new policy but does so directly by finding weights that optimize for balance between the weighted data and the target policy in the given, finite sample, which is equivalent to minimizing worst-case or posterior conditional mean square error. Our policy learner proceeds as a two-level optimization problem over policies and weights. We demonstrate that this approach markedly outperforms existing ones both in evaluation and learning, which is unsurprising given the wider support of balance-based weights. We establish extensive theoretical consistency guarantees and regret bounds that support this empirical success.

READ FULL TEXT

page 9

page 13

research
02/24/2023

Balanced Off-Policy Evaluation for Personalized Pricing

We consider a personalized pricing problem in which we have data consist...
research
08/06/2019

Policy Evaluation with Latent Confounders via Optimal Balance

Evaluating novel contextual bandit policies using logged data is crucial...
research
01/20/2023

Offline Policy Evaluation with Out-of-Sample Guarantees

We consider the problem of evaluating the performance of a decision poli...
research
03/10/2015

Doubly Robust Policy Evaluation and Optimization

We study sequential decision making in environments where rewards are on...
research
06/12/2020

Similarity-based transfer learning of decision policies

A problem of learning decision policy from past experience is considered...
research
02/19/2022

Doubly Robust Distributionally Robust Off-Policy Evaluation and Learning

Off-policy evaluation and learning (OPE/L) use offline observational dat...
research
04/13/2023

Learning Personalized Decision Support Policies

Individual human decision-makers may benefit from different forms of sup...

Please sign up or login with your details

Forgot password? Click here to reset