Analysis of the Ethiopic Twitter Dataset for Abusive Speech in Amharic

12/09/2019
by   Seid Muhie Yimam, et al.
0

In this paper, we present an analysis of the first Ethiopic Twitter Dataset for the Amharic language targeted for recognizing abusive speech. The dataset has been collected since 2014 that is written in Fidel script. Since several languages can be written using the Fidel script, we have used the existing Amharic, Tigrinya and Ge'ez corpora to retain only the Amharic tweets. We have analyzed the tweets for abusive speech content with the following targets: Analyze the distribution and tendency of abusive speech content over time and compare the abusive speech content between a Twitter and general reference Amharic corpus.

READ FULL TEXT

page 1

page 2

page 3

page 4

01/18/2022

Emojis as Anchors to Detect Arabic Offensive Language and Hate Speech

We introduce a generic, language-independent method to collect a large p...
05/12/2018

Examining a hate speech corpus for hate speech detection and popularity prediction

As research on hate speech becomes more and more relevant every day, mos...
08/05/2021

Hate Speech Detection in Roman Urdu

Hate speech is a specific type of controversial content that is widely l...
06/11/2018

Degree based Classification of Harmful Speech using Twitter Data

Harmful speech has various forms and it has been plaguing the social med...
03/13/2018

Automatic Detection of Online Jihadist Hate Speech

We have developed a system that automatically detects online jihadist ha...
12/16/2020

You Are What You Tweet: Profiling Users by Past Tweets to Improve Hate Speech Detection

Hate speech detection research has predominantly focused on purely conte...
11/27/2017

Scaling laws in geo-located Twitter data

We observe and report on a systematic relationship between population de...