Learning Mixture Model with Missing Values and its Application to Rankings

12/31/2018
by   Devavrat Shah, et al.
30

We consider the question of learning mixtures of generic sub-gaussian distributions based on observations with missing values. To that end, we utilize a matrix estimation method from literature (soft- or hard- singular value thresholding). Specifically, we stack the observations (with missing values) to form a data matrix and learn a low-rank approximation of it so that the row indices can be correctly clustered to belong to appropriate mixture component using a simple distance-based algorithm. To analyze the performance of this algorithm by quantifying finite sample bound, we extend the result for matrix estimation methods in the literature in two important ways: one, noise across columns is correlated and not independent across all entries of matrix as considered in the literature; two, the performance metric of interest is the maximum l2 row norm error, which is stronger than the traditional mean-squared-error averaged over all entries. Equipped with these advances in the context of matrix estimation, we are able to connect matrix estimation and mixture model learning in the presence of missing data.

READ FULL TEXT
research
10/26/2017

Optimal Shrinkage of Singular Values Under Random Data Contamination

A low rank matrix X has been contaminated by uniformly distributed noise...
research
02/25/2018

Time Series Analysis via Matrix Estimation

We consider the task of interpolating and forecasting a time series in t...
research
05/14/2021

Deep learned SVT: Unrolling singular value thresholding to obtain better MSE

Affine rank minimization problem is the generalized version of low rank ...
research
02/28/2019

Model Agnostic High-Dimensional Error-in-Variable Regression

We consider the problem of high-dimensional error-in-variable regression...
research
09/04/2012

Efficient EM Training of Gaussian Mixtures with Missing Data

In data-mining applications, we are frequently faced with a large fracti...
research
07/11/2022

Optimal Clustering by Lloyd Algorithm for Low-Rank Mixture Model

This paper investigates the computational and statistical limits in clus...
research
11/18/2017

Robust Synthetic Control

We present a robust generalization of the synthetic control method for c...

Please sign up or login with your details

Forgot password? Click here to reset