The Document Vectors Using Cosine Similarity Revisited

05/26/2022
by   Zhang Bingyu, et al.
0

The current state-of-the-art test accuracy (97.42%) on the IMDB movie reviews dataset was reported by <cit.> and achieved by the logistic regression classifier trained on the Document Vectors using Cosine Similarity (DV-ngrams-cosine) proposed in their paper and the Bag-of-N-grams (BON) vectors scaled by Naive Bayesian weights. While large pre-trained Transformer-based models have shown SOTA results across many datasets and tasks, the aforementioned model has not been surpassed by them, despite being much simpler and pre-trained on the IMDB dataset only. In this paper, we describe an error in the evaluation procedure of this model, which was found when we were trying to analyze its excellent performance on the IMDB dataset. We further show that the previously reported test accuracy of 97.42% is invalid and should be corrected to 93.68%. We also analyze the model performance with different amounts of training data (subsets of the IMDB dataset) and compare it to the Transformer-based RoBERTa model. The results show that while RoBERTa has a clear advantage for larger training sets, the DV-ngrams-cosine performs better than RoBERTa when the labelled training set is very small (10 or 20 documents). Finally, we introduce a sub-sampling scheme based on Naive Bayesian weights for the training process of the DV-ngrams-cosine, which leads to faster training and better quality.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/13/2022

Enhancing Few-shot Image Classification with Cosine Transformer

This paper addresses the few-shot image classification problem. One nota...
research
03/28/2022

Comparing in context: Improving cosine similarity measures with a metric tensor

Cosine similarity is a widely used measure of the relatedness of pre-tra...
research
12/27/2015

Learning Document Embeddings by Predicting N-grams for Sentiment Classification of Long Movie Reviews

Despite the loss of semantic information, bag-of-ngram based methods sti...
research
12/18/2018

Index-based, High-dimensional, Cosine Threshold Querying with Optimality Guarantees

Given a database of vectors, a cosine threshold query returns all vector...
research
02/21/2022

Embarrassingly Simple Performance Prediction for Abductive Natural Language Inference

The task of abductive natural language inference (αnli), to decide which...
research
05/11/2018

Using Stastical and Semantic Models for Multi-Document Summarization

We report on series of experiments with different semantic models on top...
research
05/10/2022

Massive Enhanced Extracted Email Features Tailored for Cosine Distance

In this paper, the process of converting the Enron email dataset (the ve...

Please sign up or login with your details

Forgot password? Click here to reset