Machine Generation and Detection of Arabic Manipulated and Fake News

by   El Moatez Billah Nagoudi, et al.

Fake news and deceptive machine-generated text are serious problems threatening modern societies, including in the Arab world. This motivates work on detecting false and manipulated stories online. However, a bottleneck for this research is lack of sufficient data to train detection models. We present a novel method for automatically generating Arabic manipulated (and potentially fake) news stories. Our method is simple and only depends on availability of true stories, which are abundant online, and a part of speech tagger (POS). To facilitate future work, we dispense with both of these requirements altogether by providing AraNews, a novel and large POS-tagged news dataset that can be used off-the-shelf. Using stories generated based on AraNews, we carry out a human annotation study that casts light on the effects of machine manipulation on text veracity. The study also measures human ability to detect Arabic machine manipulated text generated by our method. Finally, we develop the first models for detecting manipulated Arabic news and achieve state-of-the-art results on Arabic fake news detection (macro F1=70.06). Our models and data are publicly available.



There are no comments yet.


page 1

page 2

page 3

page 4


Fake or Real? A Study of Arabic Satirical Fake News

One very common type of fake news is satire which comes in a form of a n...

Detecting Cross-Modal Inconsistency to Defend Against Neural Fake News

Large-scale dissemination of disinformation online intended to mislead o...

AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset

Along with the COVID-19 pandemic, an "infodemic" of false and misleading...

Fighting Fake News: Image Splice Detection via Learned Self-Consistency

Advances in photo editing and manipulation tools have made it significan...

A Retrospective Analysis of the Fake News Challenge Stance Detection Task

The 2017 Fake News Challenge Stage 1 (FNC-1) shared task addressed a sta...

Automatic Detection of Machine Generated Text: A Critical Survey

Text generative models (TGMs) excel in producing text that matches the s...

Simple Open Stance Classification for Rumour Analysis

Stance classification determines the attitude, or stance, in a (typicall...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.