Semiparametric Estimation and Inference on Structural Target Functions using Machine Learning and Influence Functions

08/14/2020
by   Alicia Curth, et al.
15

We aim to construct a class of learning algorithms that are of practical value to applied researchers in fields such as biostatistics, epidemiology and econometrics, where the need to learn from incompletely observed information is ubiquitous. To do so, we propose a new framework for statistical machine learning, which we call 'IF-learning' due to its reliance on influence functions (IFs). To characterise the fundamental limits of what is achievable within this framework, we need to enable semiparametric estimation and inference on structural target parameters that are functions of continuous inputs arising as identifiable functionals from statistical models. Therefore, we introduce a pointwise IF to replace the true IF when it does not exist and propose learning its uncentered pointwise expected value from data. This allows us to give provable guarantees, leveraging existing general results from statistics. Our framework is problem- and model-agnostic and can be used to estimate a broad variety of target parameters of interest in applied statistics: we can consider any target function for which an IF of a population-averaged version exists in analytic form. Throughout, we put particular focus on so-called coarsening at random/doubly robust problems with partially unobserved information. This includes problems such as treatment effect estimation and inference in the presence of missing outcome data. Within this framework, we then propose two general learning algorithms that leverage ideas from the theoretical analysis: the 'IF-learner' which relies on large samples and outputs entire target functions without confidence bands, and the 'Group-IF-learner', which outputs only approximations to a function but can give confidence estimates if sufficient information on coarsening mechanisms is available. We close with a simulation study on inferring treatment effects.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/01/2022

Towards R-learner of conditional average treatment effects with a continuous treatment: T-identification, estimation, and inference

The R-learner has been popular in causal inference as a flexible and eff...
research
05/29/2022

Heterogeneous Treatment Effects Estimation: When Machine Learning meets multiple treatment regime

In many scientific and engineering domains, inferring the effect of trea...
research
05/18/2022

Non-asymptotic confidence bands on the probability an individual benefits from treatment (PIBT)

The premise of this work, in a vein similar to predictive inference with...
research
12/30/2019

Localized Debiased Machine Learning: Efficient Estimation of Quantile Treatment Effects, Conditional Value at Risk, and Beyond

We consider the efficient estimation of a low-dimensional parameter in t...
research
01/25/2019

Orthogonal Statistical Learning

We provide excess risk guarantees for statistical learning in the presen...
research
04/13/2022

Practical considerations for specifying a super learner

Common tasks encountered in epidemiology, including disease incidence es...

Please sign up or login with your details

Forgot password? Click here to reset