Tuning Traditional Language Processing Approaches for Pashto Text Classification

05/04/2023
by   Jawid Ahmad Baktash, et al.
0

Today text classification becomes critical task for concerned individuals for numerous purposes. Hence, several researches have been conducted to develop automatic text classification for national and international languages. However, the need for an automatic text categorization system for local languages is felt. The main aim of this study is to establish a Pashto automatic text classification system. In order to pursue this work, we built a Pashto corpus which is a collection of Pashto documents due to the unavailability of public datasets of Pashto text documents. Besides, this study compares several models containing both statistical and neural network machine learning techniques including Multilayer Perceptron (MLP), Support Vector Machine (SVM), K Nearest Neighbor (KNN), decision tree, gaussian naïve Bayes, multinomial naïve Bayes, random forest, and logistic regression to discover the most effective approach. Moreover, this investigation evaluates two different feature extraction methods including unigram, and Time Frequency Inverse Document Frequency (IFIDF). Subsequently, this research obtained average testing accuracy rate 94 feature extraction method in this context.

READ FULL TEXT

page 8

page 9

research
05/04/2023

Enhancing Pashto Text Classification using Language Processing Techniques for Single And Multi-Label Analysis

Text classification has become a crucial task in various fields, leading...
research
08/08/2023

A Comparative Study on TF-IDF feature Weighting Method and its Analysis using Unstructured Dataset

Text Classification is the process of categorizing text into the relevan...
research
11/15/2022

Classifying text using machine learning models and determining conversation drift

Text classification helps analyse texts for semantic meaning and relevan...
research
10/24/2018

A Text Classification Application: Poet Detection from Poetry

With the widespread use of the internet, the size of the text data incre...
research
03/15/2023

Building an Effective Email Spam Classification Model with spaCy

Today, people use email services such as Gmail, Outlook, AOL Mail, etc. ...
research
05/15/2017

Using Titles vs. Full-text as Source for Automated Semantic Document Annotation

A significant part of the largest Knowledge Graph today, the Linked Open...
research
02/01/2021

Student sentiment Analysis Using Classification With Feature Extraction Techniques

Technical growths have empowered, numerous revolutions in the educationa...

Please sign up or login with your details

Forgot password? Click here to reset