From Review to Rating: Exploring Dependency Measures for Text Classification

09/04/2017
by   Samuel Cunningham-Nelson, et al.
0

Various text analysis techniques exist, which attempt to uncover unstructured information from text. In this work, we explore using statistical dependence measures for textual classification, representing text as word vectors. Student satisfaction scores on a 3-point scale and their free text comments written about university subjects are used as the dataset. We have compared two textual representations: a frequency word representation and term frequency relationship to word vectors, and found that word vectors provide a greater accuracy. However, these word vectors have a large number of features which aggravates the burden of computational complexity. Thus, we explored using a non-linear dependency measure for feature selection by maximizing the dependence between the text reviews and corresponding scores. Our quantitative and qualitative analysis on a student satisfaction dataset shows that our approach achieves comparable accuracy to the full feature vector, while being an order of magnitude faster in testing. These text analysis and feature reduction techniques can be used for other textual data applications such as sentiment analysis.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/13/2020

Sentiment Analysis Using Averaged Weighted Word Vector Features

People use the world wide web heavily to share their experience with ent...
research
06/17/2018

An Improved Text Sentiment Classification Model Using TF-IDF and Next Word Negation

With the rapid growth of Text sentiment analysis, the demand for automat...
research
01/29/2023

Syrupy Mouthfeel and Hints of Chocolate – Predicting Coffee Review Scores using Text Based Sentiment

This paper uses textual data contained in certified (q-graded) coffee re...
research
06/08/2018

Text Classification based on Word Subspace with Term-Frequency

Text classification has become indispensable due to the rapid increase o...
research
06/14/2018

Cold-Start Aware User and Product Attention for Sentiment Classification

The use of user/product information in sentiment analysis is important, ...
research
06/12/2021

Study of sampling methods in sentiment analysis of imbalanced data

This work investigates the application of sampling methods for sentiment...
research
11/06/2017

Authorship Analysis of Xenophon's Cyropaedia

In the past several decades, many authorship attribution studies have us...

Please sign up or login with your details

Forgot password? Click here to reset