Post-Contextual-Bandit Inference

06/01/2021
by   Aurélien Bibaut, et al.
1

Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies. To support credible inference on novel interventions at the end of the study, nonetheless, we still want to construct valid confidence intervals on average treatment effects, subgroup effects, or value of new policies. The adaptive nature of the data collected by contextual bandit algorithms, however, makes this difficult: standard estimators are no longer asymptotically normally distributed and classic confidence intervals fail to provide correct coverage. While this has been addressed in non-contextual settings by using stabilized estimators, the contextual setting poses unique challenges that we tackle for the first time in this paper. We propose the Contextual Adaptive Doubly Robust (CADR) estimator, the first estimator for policy value that is asymptotically normal under contextual adaptive data collection. The main technical challenge in constructing CADR is designing adaptive and consistent conditional standard deviation estimators for stabilization. Extensive numerical experiments using 57 OpenML datasets demonstrate that confidence intervals based on CADR uniquely provide correct coverage.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/08/2020

Inference for Batched Bandits

As bandit algorithms are increasingly utilized in scientific studies, th...
research
04/29/2021

Statistical Inference with M-Estimators on Bandit Data

Bandit algorithms are increasingly used in real world sequential decisio...
research
06/07/2019

Empirical Likelihood for Contextual Bandits

We apply empirical likelihood techniques to contextual bandit policy val...
research
05/26/2023

Clip-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential Experiments

From clinical development of cancer therapies to investigations into par...
research
01/13/2023

Randomization Tests for Adaptively Collected Data

Randomization testing is a fundamental method in statistics, enabling in...
research
02/17/2023

Post-Episodic Reinforcement Learning Inference

We consider estimation and inference with data collected from episodic r...
research
02/21/2017

General Semiparametric Shared Frailty Model Estimation and Simulation with frailtySurv

The R package frailtySurv for simulating and fitting semi-parametric sha...

Please sign up or login with your details

Forgot password? Click here to reset