Investigation of Topic Modelling Methods for Understanding the Reports of the Mining Projects in Queensland

11/05/2021
by   Yasuko Okamoto, et al.
0

In the mining industry, many reports are generated in the project management process. These past documents are a great resource of knowledge for future success. However, it would be a tedious and challenging task to retrieve the necessary information if the documents are unorganized and unstructured. Document clustering is a powerful approach to cope with the problem, and many methods have been introduced in past studies. Nonetheless, there is no silver bullet that can perform the best for any types of documents. Thus, exploratory studies are required to apply the clustering methods for new datasets. In this study, we will investigate multiple topic modelling (TM) methods. The objectives are finding the appropriate approach for the mining project reports using the dataset of the Geological Survey of Queensland, Department of Resources, Queensland Government, and understanding the contents to get the idea of how to organise them. Three TM methods, Latent Dirichlet Allocation (LDA), Nonnegative Matrix Factorization (NMF), and Nonnegative Tensor Factorization (NTF) are compared statistically and qualitatively. After the evaluation, we conclude that the LDA performs the best for the dataset; however, the possibility remains that the other methods could be adopted with some improvements.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/18/2019

Topic subject creation using unsupervised learning for topic modeling

We describe the use of Non-Negative Matrix Factorization (NMF) and Laten...
research
10/22/2020

On a Guided Nonnegative Matrix Factorization

Fully unsupervised topic models have found fantastic success in document...
research
09/15/2023

Deep Nonnegative Matrix Factorization with Beta Divergences

Deep Nonnegative Matrix Factorization (deep NMF) has recently emerged as...
research
11/12/2019

Text Mining using Nonnegative Matrix Factorization and Latent Semantic Analysis

Text clustering is arguably one of the most important topics in modern d...
research
09/30/2021

A Generalized Hierarchical Nonnegative Tensor Decomposition

Nonnegative matrix factorization (NMF) has found many applications inclu...
research
04/27/2018

Can You Explain That, Better? Comprehensible Text Analytics for SE Applications

Text mining methods are used for a wide range of Software Engineering (S...
research
10/12/2021

Topic Model Supervised by Understanding Map

Inspired by the notion of Center of Mass in physics, an extension called...

Please sign up or login with your details

Forgot password? Click here to reset