Bangla Text Dataset and Exploratory Analysis for Online Harassment Detection

02/04/2021
by   Md Faisal Ahmed, et al.
0

Being the seventh most spoken language in the world, the use of the Bangla language online has increased in recent times. Hence, it has become very important to analyze Bangla text data to maintain a safe and harassment-free online place. The data that has been made accessible in this article has been gathered and marked from the comments of people in public posts by celebrities, government officials, athletes on Facebook. The total amount of collected comments is 44001. The dataset is compiled with the aim of developing the ability of machines to differentiate whether a comment is a bully expression or not with the help of Natural Language Processing and to what extent it is improper if it is an inappropriate comment. The comments are labeled with different categories of harassment. Exploratory analysis from different perspectives is also included in this paper to have a detailed overview. Due to the scarcity of data collection of categorized Bengali language comments, this dataset can have a significant role for research in detecting bully words, identifying inappropriate comments, detecting different categories of Bengali bullies, etc. The dataset is publicly available at https://data.mendeley.com/datasets/9xjx8twk8p.

READ FULL TEXT

page 1

page 2

page 3

research
09/27/2022

BanglaSarc: A Dataset for Sarcasm Detection

Being one of the most widely spoken language in the world, the use of Ba...
research
07/16/2018

LSTMs with Attention for Aggression Detection

In this paper, we describe the system submitted for the shared task on A...
research
08/21/2023

BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

Advances in automated detection of offensive language online, including ...
research
01/04/2019

Identifying Barriers to Adoption for Rust through Online Discourse

Rust is a low-level programming language known for its unique approach t...
research
07/08/2022

No Time Like the Present: Effects of Language Change on Automated Comment Moderation

The spread of online hate has become a significant problem for newspaper...
research
10/14/2020

Six Attributes of Unhealthy Conversation

We present a new dataset of approximately 44000 comments labeled by crow...
research
04/28/2020

Don't Let Me Be Misunderstood: Comparing Intentions and Perceptions in Online Discussions

Discourse involves two perspectives: a person's intention in making an u...

Please sign up or login with your details

Forgot password? Click here to reset