PELESent: Cross-domain polarity classification using distant supervision

07/09/2017
by   Edilson A. Corrêa Jr, et al.
0

The enormous amount of texts published daily by Internet users has fostered the development of methods to analyze this content in several natural language processing areas, such as sentiment analysis. The main goal of this task is to classify the polarity of a message. Even though many approaches have been proposed for sentiment analysis, some of the most successful ones rely on the availability of large annotated corpus, which is an expensive and time-consuming process. In recent years, distant supervision has been used to obtain larger datasets. So, inspired by these techniques, in this paper we extend such approaches to incorporate popular graphic symbols used in electronic messages, the emojis, in order to create a large sentiment corpus for Portuguese. Trained on almost one million tweets, several models were tested in both same domain and cross-domain corpora. Our methods obtained very competitive results in five annotated corpora from mixed domains (Twitter and product reviews), which proves the domain-independent property of such approach. In addition, our results suggest that the combination of emoticons and emojis is able to properly capture the sentiment of a message.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/15/2023

Cross-domain Sentiment Classification in Spanish

Sentiment Classification is a fundamental task in the field of Natural L...
research
08/10/2022

The Moral Foundations Reddit Corpus

Moral framing and sentiment can affect a variety of online and offline b...
research
08/01/2017

Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm

NLP tasks are often limited by scarcity of manually annotated data. In s...
research
02/06/2017

Q-WordNet PPV: Simple, Robust and (almost) Unsupervised Generation of Polarity Lexicons for Multiple Languages

This paper presents a simple, robust and (almost) unsupervised dictionar...
research
12/24/2017

Building a Sentiment Corpus of Tweets in Brazilian Portuguese

The large amount of data available in social media, forums and websites ...
research
01/23/2021

Reproducibility, Replicability and Beyond: Assessing Production Readiness of Aspect Based Sentiment Analysis in the Wild

With the exponential growth of online marketplaces and user-generated co...
research
05/17/2016

Enhanced Twitter Sentiment Classification Using Contextual Information

The rise in popularity and ubiquity of Twitter has made sentiment analys...

Please sign up or login with your details

Forgot password? Click here to reset