Regression on imperfect class labels derived by unsupervised clustering

08/16/2019
by   Rasmus Froberg Brøndum, et al.
0

Outcome regressed on class labels identified by unsupervised clustering is custom in many applications. However, it is common to ignore the misclassification of class labels caused by the learning algorithm, which potentially leads to serious bias of the estimated effect parameters. Due to its generality we suggest to redress the situation by use of the simulation and extrapolation method. Performance is illustrated by simulated data from Gaussian mixture models. Finally, we apply our method to a study which regressed overall survival on class labels derived from unsupervised clustering of gene expression data from bone marrow samples of multiple myeloma patients.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset