Abusive and Threatening Language Detection in Urdu using Supervised Machine Learning and Feature Combinations

04/06/2022
by   Muhammad Humayoun, et al.
0

This paper presents the system descriptions submitted at the FIRE Shared Task 2021 on Urdu's Abusive and Threatening Language Detection Task. This challenge aims at automatically identifying abusive and threatening tweets written in Urdu. Our submitted results were selected for the third recognition at the competition. This paper reports a non-exhaustive list of experiments that allowed us to reach the submitted results. Moreover, after the result declaration of the competition, we managed to attain even better results than the submitted results. Our models achieved 0.8318 F1 score on Task A (Abusive Language Detection for Urdu Tweets) and 0.4931 F1 score on Task B (Threatening Language Detection for Urdu Tweets). Results show that Support Vector Machines with stopwords removed, lemmatization applied, and features vector created by the combinations of word n-grams for n=1,2,3 produced the best results for Task A. For Task B, Support Vector Machines with stopwords removed, lemmatization not applied, feature vector created from a pre-trained Urdu Word2Vec (on word unigrams and bigrams), and making the dataset balanced using oversampling technique produced the best results. The code is made available for reproducibility.

READ FULL TEXT
research
04/06/2022

The 2021 Urdu Fake News Detection Task using Supervised Machine Learning and Feature Combinations

This paper presents the system description submitted at the FIRE Shared ...
research
09/22/2020

Investigating Machine Learning Methods for Language and Dialect Identification of Cuneiform Texts

Identification of the languages written using cuneiform symbols is a dif...
research
11/16/2019

General Regression Neural Networks, Radial Basis Function Neural Networks, Support Vector Machines, and Feedforward Neural Networks

The aim of this project is to develop a code to discover the optimal sig...
research
03/20/2018

UnibucKernel: A kernel-based learning method for complex word identification

In this paper, we present a kernel-based learning approach for the 2018 ...
research
06/23/2020

MANTRA: A Machine Learning reference lightcurve dataset for astronomical transient event recognition

We introduce MANTRA, an annotated dataset of 4869 transient and 71207 no...
research
10/14/2016

A Language-independent and Compositional Model for Personality Trait Recognition from Short Texts

Many methods have been used to recognize author personality traits from ...
research
03/22/2021

Identifying Machine-Paraphrased Plagiarism

Employing paraphrasing tools to conceal plagiarized text is a severe thr...

Please sign up or login with your details

Forgot password? Click here to reset