Recursive Partitioning for Heterogeneous Causal Effects

04/05/2015
by   Susan Athey, et al.
0

In this paper we study the problems of estimating heterogeneity in causal effects in experimental or observational studies and conducting inference about the magnitude of the differences in treatment effects across subsets of the population. In applications, our method provides a data-driven approach to determine which subpopulations have large or small treatment effects and to test hypotheses about the differences in these effects. For experiments, our method allows researchers to identify heterogeneity in treatment effects that was not specified in a pre-analysis plan, without concern about invalidating inference due to multiple testing. In most of the literature on supervised machine learning (e.g. regression trees, random forests, LASSO, etc.), the goal is to build a model of the relationship between a unit's attributes and an observed outcome. A prominent role in these methods is played by cross-validation which compares predictions to actual outcomes in test samples, in order to select the level of complexity of the model that provides the best predictive power. Our method is closely related, but it differs in that it is tailored for predicting causal effects of a treatment rather than a unit's outcome. The challenge is that the "ground truth" for a causal effect is not observed for any individual unit: we observe the unit with the treatment, or without the treatment, but not both at the same time. Thus, it is not obvious how to use cross-validation to determine whether a causal effect has been accurately predicted. We propose several novel cross-validation criteria for this problem and demonstrate through simulations the conditions under which they perform better than standard methods for the problem of causal effects. We then apply the method to a large-scale field experiment re-ranking results on a search engine.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/31/2017

Synth-Validation: Selecting the Best Causal Inference Method for a Given Dataset

Many decisions in healthcare, business, and other policy domains are mad...
research
07/05/2017

Machine Learning Tests for Effects on Multiple Outcomes

A core challenge in the analysis of experimental data is that the impact...
research
06/21/2022

Interpretable Deep Causal Learning for Moderation Effects

In this extended abstract paper, we address the problem of interpretabil...
research
11/17/2019

A Permutation Test for Assessing the Presence of Individual Differences in Treatment Effects

One size fits all approaches to medicine have become a thing of the past...
research
02/25/2022

Ensemble Method for Estimating Individualized Treatment Effects

In many medical and business applications, researchers are interested in...
research
06/20/2020

Learning and Testing Sub-groups with Heterogeneous Treatment Effects:A Sequence of Two Studies

There is strong interest in estimating how the magnitude of treatment ef...
research
10/16/2018

Accounting for Unobservable Heterogeneity in Cross Section Using Spatial First Differences

We propose a simple cross-sectional research design to identify causal e...

Please sign up or login with your details

Forgot password? Click here to reset