Distributed multivariable modeling for signature development under data protection constraints

03/01/2018
by   Daniela Zöller, et al.
0

Data protection constraints frequently require distributed analysis of data, i.e. individual-level data remains at many different sites, but analysis nevertheless has to be performed jointly. The data exchange is often handled manually, requiring explicit permission before transfer, i.e. the number of data calls and the amount of data should be limited. Thus, only simple summary statistics are typically transferred and aggregated with just a single call, but this does not allow for complex statistical techniques, e.g., automatic variable selection for prognostic signature development. We propose a multivariable regression approach for building a prognostic signature by automatic variable selection that is based on aggregated data from different locations in iterative calls. To minimize the amount of transferred data and the number of calls, we also provide a heuristic variant of the approach. To further strengthen data protection, the approach can also be combined with a trusted third party architecture. We evaluate our proposed method in a simulation study comparing our results to the results obtained with the pooled individual data. The proposed method is seen to be able to detect covariates with true effect to a comparable extent as a method based on individual data, although the performance is moderately decreased if the number of sites is large. In a typical scenario, the heuristic decreases the number of data calls from more than 10 to 3. To make our approach widely available for application, we provide an implementation on top of the DataSHIELD framework.

READ FULL TEXT

page 13

page 14

research
03/05/2020

Exploiting disagreement between high-dimensional variable selectors for uncertainty visualization

We propose Combined Selection and Uncertainty Visualizer (CSUV), which e...
research
09/13/2019

SuRF: a New Method for Sparse Variable Selection, with Application in Microbiome Data Analysis

In this paper, we present a new variable selection method for regression...
research
05/22/2023

Variable selection in multivariate regression model for spatially dependent data

This paper deals with variable selection in multivariate linear regressi...
research
06/28/2021

Fast Bayesian Variable Selection in Binomial and Negative Binomial Regression

Bayesian variable selection is a powerful tool for data analysis, as it ...
research
04/24/2020

Recovering individual-level spatial inference from aggregated binary data

Binary regression models are commonly used in disciplines such as epidem...
research
03/21/2022

Distributed non-disclosive validation of predictive models by a modified ROC-GLM

Distributed statistical analyses provide a promising approach for privac...
research
08/07/2018

A distributed regression analysis application based on SAS software Part II: Cox proportional hazards regression

Previous work has demonstrated the feasibility and value of conducting d...

Please sign up or login with your details

Forgot password? Click here to reset