A Simple and Efficient Ensemble Classifier Combining Multiple Neural Network Models on Social Media Datasets in Vietnamese

09/28/2020
by   Huy Duc Huynh, et al.
0

Text classification is a popular topic of natural language processing, which has currently attracted numerous research efforts worldwide. The significant increase of data in social media requires the vast attention of researchers to analyze such data. There are various studies in this field in many languages but limited to the Vietnamese language. Therefore, this study aims to classify Vietnamese texts on social media from three different Vietnamese benchmark datasets. Advanced deep learning models are used and optimized in this study, including CNN, LSTM, and their variants. We also implement the BERT, which has never been applied to the datasets. Our experiments find a suitable model for classification tasks on each specific dataset. To take advantage of single models, we propose an ensemble model, combining the highest-performance models. Our single models reach positive results on each dataset. Moreover, our ensemble model achieves the best performance on all three datasets. We reach 86.96 UIT-VSMEC dataset, 92.79 dataset, respectively. Therefore, our models achieve better performances as compared to previous studies on these datasets.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/21/2022

SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese

Text classification is a typical natural language processing or computat...
research
11/11/2016

UTCNN: a Deep Learning Model of Stance Classificationon on Social Media Text

Most neural network models for document classification on social media f...
research
08/31/2018

Neural DrugNet

In this paper, we describe the system submitted for the shared task on S...
research
07/09/2021

A Robust Deep Ensemble Classifier for Figurative Language Detection

Recognition and classification of Figurative Language (FL) is an open pr...
research
10/31/2020

Rumor Detection on Twitter Using Multiloss Hierarchical BiLSTM with an Attenuation Factor

Social media platforms such as Twitter have become a breeding ground for...
research
06/30/2023

A Cost-aware Study of Depression Language on Social Media using Topic and Affect Contextualization

Depression is a growing issue in society's mental health that affects al...
research
12/30/2022

How would Stance Detection Techniques Evolve after the Launch of ChatGPT?

Stance detection refers to the task of extracting the standpoint (Favor,...

Please sign up or login with your details

Forgot password? Click here to reset