Leveraging text data for causal inference using electronic health records

06/09/2023
by   Reagan Mozer, et al.
0

Text is a ubiquitous component of medical data, containing valuable information about patient characteristics and care that are often missing from structured chart data. Despite this richness, it is rarely used in clinical research, owing partly to its complexity. Using a large database of patient records and treatment histories accompanied by extensive notes by attendant physicians and nurses, we show how text data can be used to support causal inference with electronic health data in all stages, from conception and design to analysis and interpretation, with minimal additional effort. We focus on studies using matching for causal inference. We augment a classic matching analysis by incorporating text in three ways: by using text to supplement a multiple imputation procedure, we improve the fidelity of imputed values to handle missing data; by incorporating text in the matching stage, we strengthen the plausibility of the matching procedure; and by conditioning on text, we can estimate easily interpretable text-based heterogeneous treatment effects that may be stronger than those found across categories of structured covariates. Using these techniques, we hope to expand the scope of secondary analysis of clinical data to domains where quantitative data is of poor quality or nonexistent, but where text is available, such as in developing countries.

READ FULL TEXT
research
02/25/2020

MissDeepCausal: Causal Inference from Incomplete Data Using Deep Latent Variable Models

Inferring causal effects of a treatment, intervention or policy from obs...
research
07/09/2021

Hypothetical estimands in clinical trials: a unification of causal inference and missing data methods

The ICH E9 addendum introduces the term intercurrent event to refer to e...
research
05/06/2019

Estimating the effect of PEG in ALS patients using observational data subject to censoring by death and missing outcomes

Though they may offer valuable patient and disease information that is i...
research
11/26/2019

Hybrid Text Feature Modeling for Disease Group Prediction using Unstructured Physician Notes

Existing Clinical Decision Support Systems (CDSSs) largely depend on the...
research
08/10/2020

Using Multiple Imputation to Classify Potential Outcomes Subgroups

With medical tests becoming increasingly available, concerns about over-...
research
01/13/2019

Propensity scores using missingness pattern information: a practical guide

Electronic health records are a valuable data source for investigating h...
research
11/11/2020

Teaching deep learning causal effects improves predictive performance

Causal inference is a powerful statistical methodology for explanatory a...

Please sign up or login with your details

Forgot password? Click here to reset