Empirical Bayes PCA in high dimensions

by   Xinyi Zhong, et al.

When the dimension of data is comparable to or larger than the number of available data samples, Principal Components Analysis (PCA) is known to exhibit problematic phenomena of high-dimensional noise. In this work, we propose an Empirical Bayes PCA method that reduces this noise by estimating a structural prior for the joint distributions of the principal components. This EB-PCA method is based upon the classical Kiefer-Wolfowitz nonparametric MLE for empirical Bayes estimation, distributional results derived from random matrix theory for the sample PCs, and iterative refinement using an Approximate Message Passing (AMP) algorithm. In theoretical "spiked" models, EB-PCA achieves Bayes-optimal estimation accuracy in the same settings as the oracle Bayes AMP procedure that knows the true priors. Empirically, EB-PCA can substantially improve over PCA when there is strong prior structure, both in simulation and on several quantitative benchmarks constructed using data from the 1000 Genomes Project and the International HapMap Project. A final illustration is presented for an analysis of gene expression data obtained by single-cell RNA-seq.


PCA Initialization for Approximate Message Passing in Rotationally Invariant Models

We study the problem of estimating a rank-1 signal in the presence of ro...

A nonparametric empirical Bayes approach to covariance matrix estimation

We propose an empirical Bayes method to estimate high-dimensional covari...

Empirical Bayes Matrix Factorization

Matrix factorization methods - including Factor analysis (FA), and Princ...

Selecting the number of components in PCA via random signflips

Dimensionality reduction via PCA and factor analysis is an important too...

Approximate Message Passing algorithms for rotationally invariant matrices

Approximate Message Passing (AMP) algorithms have seen widespread use ac...

Data Distillery: Effective Dimension Estimation via Penalized Probabilistic PCA

The paper tackles the unsupervised estimation of the effective dimension...

Comparison of Canonical Correlation and Partial Least Squares analyses of simulated and empirical data

In this paper, we compared the general forms of CCA and PLS on three sim...