Refined Complexity of PCA with Outliers

05/10/2019
by   Fedor V. Fomin, et al.
0

Principal component analysis (PCA) is one of the most fundamental procedures in exploratory data analysis and is the basic step in applications ranging from quantitative finance and bioinformatics to image analysis and neuroscience. However, it is well-documented that the applicability of PCA in many real scenarios could be constrained by an "immune deficiency" to outliers such as corrupted observations. We consider the following algorithmic question about the PCA with outliers. For a set of n points in R^d, how to learn a subset of points, say 1 remaining part of the points is best fit into some unknown r-dimensional subspace? We provide a rigorous algorithmic analysis of the problem. We show that the problem is solvable in time n^O(d^2). In particular, for constant dimension the problem is solvable in polynomial time. We complement the algorithmic result by the lower bound, showing that unless Exponential Time Hypothesis fails, in time f(d)n^o(d), for any function f of d, it is impossible not only to solve the problem exactly but even to approximate it within a constant factor.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/02/2012

Robust Principal Component Analysis Using Statistical Estimators

Principal Component Analysis (PCA) finds a linear mapping and maximizes ...
research
08/07/2020

Modal Principal Component Analysis

Principal component analysis (PCA) is a widely used method for data proc...
research
03/31/2021

On the Optimality of the Oja's Algorithm for Online PCA

In this paper we analyze the behavior of the Oja's algorithm for online/...
research
11/29/2019

Adversarially Robust Low Dimensional Representations

Adversarial or test time robustness measures the susceptibility of a mac...
research
09/15/2016

Coherence Pursuit: Fast, Simple, and Robust Principal Component Analysis

This paper presents a remarkably simple, yet powerful, algorithm termed ...
research
06/17/2021

Pre-treatment of outliers and anomalies in plant data: Methodology and case study of a Vacuum Distillation Unit

Data pre-treatment plays a significant role in improving data quality, t...
research
05/24/2020

Derivation of Symmetric PCA Learning Rules from a Novel Objective Function

Neural learning rules for principal component / subspace analysis (PCA /...

Please sign up or login with your details

Forgot password? Click here to reset