KeyXtract Twitter Model - An Essential Keywords Extraction Model for Twitter Designed using NLP Tools

08/09/2017
by   Tharindu Weerasooriya, et al.
0

Since a tweet is limited to 140 characters, it is ambiguous and difficult for traditional Natural Language Processing (NLP) tools to analyse. This research presents KeyXtract which enhances the machine learning based Stanford CoreNLP Part-of-Speech (POS) tagger with the Twitter model to extract essential keywords from a tweet. The system was developed using rule-based parsers and two corpora. The data for the research was obtained from a Twitter profile of a telecommunication company. The system development consisted of two stages. At the initial stage, a domain specific corpus was compiled after analysing the tweets. The POS tagger extracted the Noun Phrases and Verb Phrases while the parsers removed noise and extracted any other keywords missed by the POS tagger. The system was evaluated using the Turing Test. After it was tested and compared against Stanford CoreNLP, the second stage of the system was developed addressing the shortcomings of the first stage. It was enhanced using Named Entity Recognition and Lemmatization. The second stage was also tested using the Turing test and its pass rate increased from 50.00 performance of the final system output was measured using the F1 score. Stanford CoreNLP with the Twitter model had an average F1 of 0.69 while the improved system had a F1 of 0.77. The accuracy of the system could be improved by using a complete domain specific corpus. Since the system used linguistic features of a sentence, it could be applied to other NLP tools.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/01/2020

Detecting White Supremacist Hate Speech using Domain Specific Word Embedding with Deep Learning and BERT

White supremacists embrace a radical ideology that considers white peopl...
research
01/01/2023

Floods Relevancy and Identification of Location from Twitter Posts using NLP Techniques

This paper presents our solutions for the MediaEval 2022 task on Disaste...
research
06/22/2023

Named entity recognition in resumes

Named entity recognition (NER) is used to extract information from vario...
research
10/22/2020

Method of noun phrase detection in Ukrainian texts

Introduction. The area of natural language processing considers AI-compl...
research
09/15/2020

Improving Joint Layer RNN based Keyphrase Extraction by Using Syntactical Features

Keyphrase extraction as a task to identify important words or phrases fr...
research
03/23/2022

Multi-Mosaics: Corpus Summarizing and Exploration using multiple Concordance Mosaic Visualisations

Researchers working in areas such as lexicography, translation studies, ...
research
02/22/2017

Dialectometric analysis of language variation in Twitter

In the last few years, microblogging platforms such as Twitter have give...

Please sign up or login with your details

Forgot password? Click here to reset