Uncovering Political Hate Speech During Indian Election Campaign: A New Low-Resource Dataset and Baselines

06/26/2023
by   Farhan Ahmad Jafri, et al.
0

The detection of hate speech in political discourse is a critical issue, and this becomes even more challenging in low-resource languages. To address this issue, we introduce a new dataset named IEHate, which contains 11,457 manually annotated Hindi tweets related to the Indian Assembly Election Campaign from November 1, 2021, to March 9, 2022. We performed a detailed analysis of the dataset, focusing on the prevalence of hate speech in political communication and the different forms of hateful language used. Additionally, we benchmark the dataset using a range of machine learning, deep learning, and transformer-based algorithms. Our experiments reveal that the performance of these models can be further improved, highlighting the need for more advanced techniques for hate speech detection in low-resource languages. In particular, the relatively higher score of human evaluation over algorithms emphasizes the importance of utilizing both human and automated approaches for effective hate speech moderation. Our IEHate dataset can serve as a valuable resource for researchers and practitioners working on developing and evaluating hate speech detection techniques in low-resource languages. Overall, our work underscores the importance of addressing the challenges of identifying and mitigating hate speech in political discourse, particularly in the context of low-resource languages. The dataset and resources for this work are made available at https://github.com/Farhan-jafri/Indian-Election.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/14/2020

Deep Learning Models for Multilingual Hate Speech Detection

Hate speech detection is a challenging problem with most of the datasets...
research
09/06/2023

RoDia: A New Dataset for Romanian Dialect Identification from Speech

Dialect identification is a critical task in speech processing and langu...
research
03/29/2023

Tackling Hate Speech in Low-resource Languages with Context Experts

Given Myanmars historical and socio-political context, hate speech sprea...
research
06/30/2019

Evaluating Language Model Finetuning Techniques for Low-resource Languages

Unlike mainstream languages (such as English and French), low-resource l...
research
03/13/2020

LSCP: Enhanced Large Scale Colloquial Persian Language Understanding

Language recognition has been significantly advanced in recent years by ...
research
10/23/2022

A Greek Parliament Proceedings Dataset for Computational Linguistics and Political Analysis

Large, diachronic datasets of political discourse are hard to come acros...
research
05/03/2022

BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions

Parliamentary transcripts provide a valuable resource to understand the ...

Please sign up or login with your details

Forgot password? Click here to reset