Cyber Hate Classification: 'Othering' Language And Paragraph Embedding

01/23/2018
by   Wafa Alorainy, et al.
0

Hateful and offensive language (also known as hate speech or cyber hate) posted and widely circulated via the World Wide Web can be considered as a key risk factor for individual and societal tension linked to regional instability. Automated Web-based hate speech detection is important for the observation and understanding trends of societal tension. In this research, we improve on existing research by proposing different data mining feature extraction methods. While previous work has involved using lexicons, bags-of-words or probabilistic parsing approach (e.g. using Typed Dependencies), they all suffer from a similar issue which is that hate speech can often be subtle and indirect, and depending on individual words or phrases can lead to a significant number of false negatives. This problem motivated us to conduct new experiments to identify subtle language use, such as references to immigration or job prosperity in a hateful context. We propose a novel 'Othering Lexicon' to identify these subtleties and we incorporate our lexicon with embedding learning for feature extraction and subsequent classification using a neural network approach. Our method first explores the context around othering terms in a corpus, and identifies context patterns that are relevant to the othering context. These patterns are used along with the othering pronoun and hate speech terms to build our 'Othering Lexicon'. Embedding algorithm has the superior characteristic that the similar words have a closer distance, which is helpful to train our classifier on the negative and positive classes. For validation, several experiments were conducted on different types of hate speech, namely religion, disability, race and sexual orientation, with F-measure scores for classifying hateful instances obtained through applying our model of 0.93, 0.95, 0.97 and 0.92 respective.

READ FULL TEXT
research
01/23/2018

The Enemy Among Us: Detecting Hate Speech with Threats Based 'Othering' Language Embeddings

Offensive or antagonistic language targeted at individuals and social gr...
research
10/20/2017

Detecting Online Hate Speech Using Context Aware Models

In the wake of a polarizing election, the cyber world is laden with hate...
research
07/07/2022

Multimodal Feature Extraction for Memes Sentiment Classification

In this study, we propose feature extraction for multimodal meme classif...
research
08/27/2021

Speech Representations and Phoneme Classification for Preserving the Endangered Language of Ladin

A vast majority of the world's 7,000 spoken languages are predicted to b...
research
04/27/2017

A Survey of Neural Network Techniques for Feature Extraction from Text

This paper aims to catalyze the discussions about text feature extractio...
research
02/07/2018

Classification of Things in DBpedia using Deep Neural Networks

The Semantic Web aims at representing knowledge about the real world at ...

Please sign up or login with your details

Forgot password? Click here to reset