A new LDA formulation with covariates

02/18/2022
by   Gilson Shimizu, et al.
0

The Latent Dirichlet Allocation (LDA) model is a popular method for creating mixed-membership clusters. Despite having been originally developed for text analysis, LDA has been used for a wide range of other applications. We propose a new formulation for the LDA model which incorporates covariates. In this model, a negative binomial regression is embedded within LDA, enabling straight-forward interpretation of the regression coefficients and the analysis of the quantity of cluster-specific elements in each sampling units (instead of the analysis being focused on modeling the proportion of each cluster, as in Structural Topic Models). We use slice sampling within a Gibbs sampling algorithm to estimate model parameters. We rely on simulations to show how our algorithm is able to successfully retrieve the true parameter values and the ability to make predictions for the abundance matrix using the information given by the covariates. The model is illustrated using real data sets from three different areas: text-mining of Coronavirus articles, analysis of grocery shopping baskets, and ecology of tree species on Barro Colorado Island (Panama). This model allows the identification of mixed-membership clusters in discrete data and provides inference on the relationship between covariates and the abundance of these clusters.

READ FULL TEXT

page 19

page 25

research
09/12/2016

Hyperspectral Unmixing with Endmember Variability using Partial Membership Latent Dirichlet Allocation

The application of Partial Membership Latent Dirichlet Allocation(PM-LDA...
research
08/24/2018

Measuring LDA Topic Stability from Clusters of Replicated Runs

Background: Unstructured and textual data is increasing rapidly and Late...
research
04/21/2020

Revealing Cluster Structures Based on Mixed Sampling Frequencies

This paper proposes a new nonparametric mixed data sampling (MIDAS) mode...
research
10/09/2020

Latent Dirichlet Allocation Model Training with Differential Privacy

Latent Dirichlet Allocation (LDA) is a popular topic modeling technique ...
research
08/29/2016

What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)

Context: Topic modeling finds human-readable structures in unstructured ...
research
09/11/2021

Microbiome subcommunity learning with logistic-tree normal latent Dirichlet allocation

Mixed-membership (MM) models such as Latent Dirichlet Allocation (LDA) h...
research
10/01/2021

ALBU: An approximate Loopy Belief message passing algorithm for LDA to improve performance on small data sets

Variational Bayes (VB) applied to latent Dirichlet allocation (LDA) has ...

Please sign up or login with your details

Forgot password? Click here to reset