Robust self-tuning semiparametric PCA for contaminated elliptical distribution

06/08/2022
by   Hung Hung, et al.
0

Principal component analysis (PCA) is one of the most popular dimension reduction methods. The usual PCA is known to be sensitive to the presence of outliers, and thus many robust PCA methods have been developed. Among them, the Tyler's M-estimator is shown to be the most robust scatter estimator under the elliptical distribution. However, when the underlying distribution is contaminated and deviates from ellipticity, Tyler's M-estimator might not work well. In this article, we apply the semiparametric theory to propose a robust semiparametric PCA. The merits of our proposal are twofold. First, it is robust to heavy-tailed elliptical distributions as well as robust to non-elliptical outliers. Second, it pairs well with a data-driven tuning procedure, which is based on active ratio and can adapt to different degrees of data outlyingness. Theoretical properties are derived, including the influence functions for various statistical functionals and asymptotic normality. Simulation studies and a data analysis demonstrate the superiority of our method.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/07/2020

Modal Principal Component Analysis

Principal component analysis (PCA) is a widely used method for data proc...
research
02/22/2023

On the efficiency-loss free ordering-robustness of product-PCA

This article studies the robustness of the eigenvalue ordering, an impor...
research
11/25/2022

Data-driven identification and analysis of the glass transition in polymer melts

We propose a data-driven approach based on information about structural ...
research
06/17/2021

Pre-treatment of outliers and anomalies in plant data: Methodology and case study of a Vacuum Distillation Unit

Data pre-treatment plays a significant role in improving data quality, t...
research
05/12/2023

Robust score matching for compositional data

The restricted polynomially-tilted pairwise interaction (RPPI) distribut...
research
04/29/2022

Distributed Learning for Principle Eigenspaces without Moment Constraints

Distributed Principal Component Analysis (PCA) has been studied to deal ...
research
06/04/2018

MacroPCA: An all-in-one PCA method allowing for missing values as well as cellwise and rowwise outliers

Multivariate data are typically represented by a rectangular matrix (tab...

Please sign up or login with your details

Forgot password? Click here to reset