A Tensor-EM Method for Large-Scale Latent Class Analysis with Clustering Consistency

03/30/2021
by   Zhenghao Zeng, et al.
0

Latent class models are powerful statistical modeling tools widely used in psychological, behavioral, and social sciences. In the modern era of data science, researchers often have access to response data collected from large-scale surveys or assessments, featuring many items (large J) and many subjects (large N). This is in contrary to the traditional regime with fixed J and large N. To analyze such large-scale data, it is important to develop methods that are both computationally efficient and theoretically valid. In terms of computation, the conventional EM algorithm for latent class models tends to have a slow algorithmic convergence rate for large-scale data and may converge to some local optima instead of the maximum likelihood estimator (MLE). Motivated by this, we introduce the tensor decomposition perspective into latent class analysis. Methodologically, we propose to use a moment-based tensor power method in the first step, and then use the obtained estimators as initialization for the EM algorithm in the second step. Theoretically, we establish the clustering consistency of the MLE in assigning subjects into latent classes when N and J both go to infinity. Simulation studies suggest that the proposed tensor-EM pipeline enjoys both good accuracy and computational efficiency for large-scale data. We also apply the proposed method to a personality dataset as an illustration.

READ FULL TEXT
research
01/09/2019

Beyond the EM Algorithm: Constrained Optimization Methods for Latent Class Model

Latent class model (LCM), which is a finite mixture of different categor...
research
12/19/2017

Joint Maximum Likelihood Estimation for High-dimensional Exploratory Item Response Analysis

Multidimensional item response theory is widely used in education and ps...
research
12/24/2017

Structured Latent Factor Analysis for Large-scale Data: Identifiability, Estimability, and Their Implications

Latent factor models are widely used to measure unobserved latent traits...
research
08/29/2020

Statistical Analysis of Multi-Relational Network Recovery

In this paper, we develop asymptotic theories for a class of latent vari...
research
07/04/2023

Local primordial non-Gaussianity from the large-scale clustering of photometric DESI luminous red galaxies

We use angular clustering of luminous red galaxies from the Dark Energy ...
research
06/21/2023

High-dimensional Tensor Response Regression using the t-Distribution

In recent years, promising statistical modeling approaches to tensor dat...
research
08/29/2020

Subtask Analysis of Process Data Through a Predictive Model

Response process data collected from human-computer interactive items co...

Please sign up or login with your details

Forgot password? Click here to reset