Text Classification Using Hybrid Machine Learning Algorithms on Big Data

03/30/2021
by   D. C. Asogwa, et al.
0

Recently, there are unprecedented data growth originating from different online platforms which contribute to big data in terms of volume, velocity, variety and veracity (4Vs). Given this nature of big data which is unstructured, performing analytics to extract meaningful information is currently a great challenge to big data analytics. Collecting and analyzing unstructured textual data allows decision makers to study the escalation of comments/posts on our social media platforms. Hence, there is need for automatic big data analysis to overcome the noise and the non-reliability of these unstructured dataset from the digital media platforms. However, current machine learning algorithms used are performance driven focusing on the classification/prediction accuracy based on known properties learned from the training samples. With the learning task in a large dataset, most machine learning models are known to require high computational cost which eventually leads to computational complexity. In this work, two supervised machine learning algorithms are combined with text mining techniques to produce a hybrid model which consists of Naïve Bayes and support vector machines (SVM). This is to increase the efficiency and accuracy of the results obtained and also to reduce the computational cost and complexity. The system also provides an open platform where a group of persons with a common interest can share their comments/messages and these comments classified automatically as legal or illegal. This improves the quality of conversation among users. The hybrid model was developed using WEKA tools and Java programming language. The result shows that the hybrid model gave 96.76 69.21

READ FULL TEXT
research
03/18/2015

Efficient Machine Learning for Big Data: A Review

With the emerging technologies and all associated devices, it is predict...
research
07/23/2014

Using 3D Printing to Visualize Social Media Big Data

Big data volume continues to grow at unprecedented rates. One of the key...
research
04/17/2018

Rafiki: Machine Learning as an Analytics Service System

Big data analytics is gaining massive momentum in the last few years. Ap...
research
09/21/2022

Benchmarking Apache Spark and Hadoop MapReduce on Big Data Classification

Most of the popular Big Data analytics tools evolved to adapt their work...
research
09/23/2022

KeypartX: Graph-based Perception (Text) Representation

The availability of big data has opened up big opportunities for individ...
research
11/12/2020

Occams Razor for Big Data? On Detecting Quality in Large Unstructured Datasets

Detecting quality in large unstructured datasets requires capacities far...
research
02/10/2018

Document Classification Using Distributed Machine Learning

In this paper, we investigate the performance and success rates of Naïve...

Please sign up or login with your details

Forgot password? Click here to reset