Learning Topics using Semantic Locality

04/11/2018
by   Ziyi Zhao, et al.
0

The topic modeling discovers the latent topic probability of the given text documents. To generate the more meaningful topic that better represents the given document, we proposed a new feature extraction technique which can be used in the data preprocessing stage. The method consists of three steps. First, it generates the word/word-pair from every single document. Second, it applies a two-way TF-IDF algorithm to word/word-pair for semantic filtering. Third, it uses the K-means algorithm to merge the word pairs that have the similar semantic meaning. Experiments are carried out on the Open Movie Database (OMDb), Reuters Dataset and 20NewsGroup Dataset. The mean Average Precision score is used as the evaluation metric. Comparing our results with other state-of-the-art topic models, such as Latent Dirichlet allocation and traditional Restricted Boltzmann Machines. Our proposed data preprocessing can improve the generated topic accuracy by up to 12.99%.

READ FULL TEXT

page 2

page 5

page 6

research
01/25/2023

Improving the Inference of Topic Models via Infinite Latent State Replications

In text mining, topic models are a type of probabilistic generative mode...
research
04/26/2020

Neural Topic Modeling with Bidirectional Adversarial Training

Recent years have witnessed a surge of interests of using neural topic m...
research
04/05/2016

Feature extraction using Latent Dirichlet Allocation and Neural Networks: A case study on movie synopses

Feature extraction has gained increasing attention in the field of machi...
research
03/09/2019

A New Approach for Topic Detection using Adaptive Neural Networks

Topic detection becomes more important due to the increase of informatio...
research
05/01/2016

Text-mining the NeuroSynth corpus using Deep Boltzmann Machines

Large-scale automated meta-analysis of neuroimaging data has recently es...
research
09/12/2018

Semantic WordRank: Generating Finer Single-Document Summarizations

We present Semantic WordRank (SWR), an unsupervised method for generatin...
research
07/10/2020

Handling Collocations in Hierarchical Latent Tree Analysis for Topic Modeling

Topic modeling has been one of the most active research areas in machine...

Please sign up or login with your details

Forgot password? Click here to reset