Comparative Performance of Machine Learning Algorithms in Cyberbullying Detection: Using Turkish Language Preprocessing Techniques

01/29/2021
by   Emre Cihan Ates, et al.
0

With the increasing use of the internet and social media, it is obvious that cyberbullying has become a major problem. The most basic way for protection against the dangerous consequences of cyberbullying is to actively detect and control the contents containing cyberbullying. When we look at today's internet and social media statistics, it is impossible to detect cyberbullying contents only by human power. Effective cyberbullying detection methods are necessary in order to make social media a safe communication space. Current research efforts focus on using machine learning for detecting and eliminating cyberbullying. Although most of the studies have been conducted on English texts for the detection of cyberbullying, there are few studies in Turkish. Limited methods and algorithms were also used in studies conducted on the Turkish language. In addition, the scope and performance of the algorithms used to classify the texts containing cyberbullying is different, and this reveals the importance of using an appropriate algorithm. The aim of this study is to compare the performance of different machine learning algorithms in detecting Turkish messages containing cyberbullying. In this study, nineteen different classification algorithms were used to identify texts containing cyberbullying using Turkish natural language processing techniques. Precision, recall, accuracy and F1 score values were used to evaluate the performance of classifiers. It was determined that the Light Gradient Boosting Model (LGBM) algorithm showed the best performance with 90.788 Score value.

READ FULL TEXT

page 7

page 10

page 11

page 13

page 14

page 15

research
01/25/2022

Suicidal Ideation Detection on Social Media: A Review of Machine Learning Methods

Social media platforms have transformed traditional communication method...
research
01/07/2021

Detecting Suspicious Events in Fast Information Flows

We describe a computational feather-light and intuitive, yet provably ef...
research
04/02/2019

Effectiveness of Data-Driven Induction of Semantic Spaces and Traditional Classifiers for Sarcasm Detection

Irony and sarcasm are two complex linguistic phenomena that are widely u...
research
08/29/2023

Vulgar Remarks Detection in Chittagonian Dialect of Bangla

The negative effects of online bullying and harassment are increasing wi...
research
12/29/2017

Methods for Detecting Paraphrase Plagiarism

Paraphrase plagiarism is one of the difficult challenges facing plagiari...
research
06/01/2022

Vietnamese Hate and Offensive Detection using PhoBERT-CNN and Social Media Streaming Data

Society needs to develop a system to detect hate and offense to build a ...
research
12/03/2019

See and Read: Detecting Depression Symptoms in Higher Education Students Using Multimodal Social Media Data

Mental disorders such as depression and anxiety have been increasing at ...

Please sign up or login with your details

Forgot password? Click here to reset