Semi-Counterfactual Risk Minimization Via Neural Networks

09/15/2022
by   Gholamali Aminian, et al.
4

Counterfactual risk minimization is a framework for offline policy optimization with logged data which consists of context, action, propensity score, and reward for each sample point. In this work, we build on this framework and propose a learning method for settings where the rewards for some samples are not observed, and so the logged data consists of a subset of samples with unknown rewards and a subset of samples with known rewards. This setting arises in many application domains, including advertising and healthcare. While reward feedback is missing for some samples, it is possible to leverage the unknown-reward samples in order to minimize the risk, and we refer to this setting as semi-counterfactual risk minimization. To approach this kind of learning problem, we derive new upper bounds on the true risk under the inverse propensity score estimator. We then build upon these bounds to propose a regularized counterfactual risk minimization method, where the regularization term is based on the logged unknown-rewards dataset only; hence it is reward-independent. We also propose another algorithm based on generating pseudo-rewards for the logged unknown-rewards dataset. Experimental results with neural networks and benchmark datasets indicate that these algorithms can leverage the logged unknown-rewards dataset besides the logged known-reward dataset.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/29/2018

Bayesian Counterfactual Risk Minimization

We present a Bayesian view of counterfactual risk minimization (CRM), al...
research
02/23/2023

Sequential Counterfactual Risk Minimization

Counterfactual Risk Minimization (CRM) is a framework for dealing with t...
research
04/22/2020

Optimization Approaches for Counterfactual Risk Minimization with Continuous Actions

Counterfactual reasoning from logged data has become increasingly import...
research
12/23/2016

Constructing Effective Personalized Policies Using Counterfactual Inference from Biased Data Sets with Many Features

This paper proposes a novel approach for constructing effective personal...
research
06/14/2019

Distributionally Robust Counterfactual Risk Minimization

This manuscript introduces the idea of using Distributionally Robust Opt...
research
02/09/2015

Counterfactual Risk Minimization: Learning from Logged Bandit Feedback

We develop a learning principle and an efficient algorithm for batch lea...
research
07/25/2020

Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions

Users of music streaming, video streaming, news recommendation, and e-co...

Please sign up or login with your details

Forgot password? Click here to reset