A Benchmark Dataset for Learning to Intervene in Online Hate Speech

by   Jing Qian, et al.

Countering online hate speech is a critical yet challenging task, but one which can be aided by the use of Natural Language Processing (NLP) techniques. Previous research has primarily focused on the development of NLP methods to automatically and effectively detect online hate speech while disregarding further action needed to calm and discourage individuals from using hate speech in the future. In addition, most existing hate speech datasets treat each post as an isolated instance, ignoring the conversational context. In this paper, we propose a novel task of generative hate speech intervention, where the goal is to automatically generate responses to intervene during online conversations that contain hate speech. As a part of this work, we introduce two fully-labeled large-scale hate speech intervention datasets collected from Gab and Reddit. These datasets provide conversation segments, hate speech labels, as well as intervention responses written by Mechanical Turk Workers. In this paper, we also analyze the datasets to understand the common intervention strategies and explore the performance of common automatic response generation methods on these new datasets to provide a benchmark for future research.


page 1

page 2

page 3

page 4


Countering Online Hate Speech: An NLP Perspective

Online hate speech has caught everyone's attention from the news related...

Using Synthetic Data for Conversational Response Generation in Low-resource Settings

Response generation is a task in natural language processing (NLP) where...

Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech

Countermeasures to effectively fight the ever increasing hate speech onl...

Towards generalisable hate speech detection: a review on obstacles and solutions

Hate speech is one type of harmful online content which directly attacks...

Automated Utterance Labeling of Conversations Using Natural Language Processing

Conversational data is essential in psychology because it can help resea...

Sarcasm Detection using Context Separators in Online Discourse

Sarcasm is an intricate form of speech, where meaning is conveyed implic...

Evaluation of In-Person Counseling Strategies To Develop Physical Activity Chatbot for Women

Artificial intelligence chatbots are the vanguard in technology-based in...

Please sign up or login with your details

Forgot password? Click here to reset