Using permutations to quantify and correct for confounding in machine learning predictions

05/18/2018
by   Elias Chaibub Neto, et al.
0

Clinical machine learning applications are often plagued with confounders that are clinically irrelevant, but can still artificially boost the predictive performance of the algorithms. Confounding is especially problematic in mobile health studies run "in the wild", where it is challenging to balance the demographic characteristics of participants that self select to enter the study. Here, we develop novel permutation approaches to quantify and adjust for the influence of observed confounders in machine learning predictions. Using restricted permutations we develop statistical tests to detect response learning in the presence of confounding, as well as, confounding learning per se. In particular, we prove that restricted permutations provide an alternative method to compute partial correlations. This result motivates a novel approach to adjust for confounders, where we are able to "subtract" the contribution of the confounders from the observed predictive performance of a machine learning algorithm using a mapping between restricted and standard permutation null distributions. We evaluate the statistical properties of our approach in simulation studies, and illustrate its application to synthetic data sets.

READ FULL TEXT
research
11/29/2018

Using permutations to assess confounding in machine learning applications for digital health

Clinical machine learning applications are often plagued with confounder...
research
11/01/2021

Statistical quantification of confounding bias in predictive modelling

The lack of non-parametric statistical tests for confounding bias signif...
research
12/08/2017

Detecting confounding due to subject identification in clinical machine learning diagnostic applications: a permutation test approach

Recently, Saeb et al (2017) showed that, in diagnostic machine learning ...
research
12/19/2022

Counterfactual Risk Assessments under Unmeasured Confounding

Statistical risk assessments inform consequential decisions, such as pre...
research
10/27/2022

A Double Machine Learning Trend Model for Citizen Science Data

1. Citizen and community-science (CS) datasets have great potential for ...

Please sign up or login with your details

Forgot password? Click here to reset