A Framework for Fast Polarity Labelling of Massive Data Streams

03/23/2022
by   Huilin Wu, et al.
0

Many of the existing sentiment analysis techniques are based on supervised learning, and they demand the availability of valuable training datasets to train their models. When dataset freshness is critical, the annotating of high speed unlabelled data streams becomes critical but remains an open problem. In this paper, we propose PLStream, a novel Apache Flink-based framework for fast polarity labelling of massive data streams, like Twitter tweets or online product reviews. We address the associated implementation challenges and propose a list of techniques including both algorithmic improvements and system optimizations. A thorough empirical validation with two real-world workloads demonstrates that PLStream is able to generate high quality labels (almost 80 accuracy) in the presence of high-speed continuous unlabelled data streams (almost 16,000 tuples/sec) without any manual efforts.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/05/2022

A Comparison of Automatic Labelling Approaches for Sentiment Analysis

Labelling a large quantity of social media data for the task of supervis...
research
09/25/2015

Sentiment Uncertainty and Spam in Twitter Streams and Its Implications for General Purpose Realtime Sentiment Analysis

State of the art benchmarks for Twitter Sentiment Analysis do not consid...
research
12/16/2019

A new Frequency Estimation Sketch for Data Streams

In data stream applications, one of the critical issues is to estimate t...
research
03/31/2023

A Benchmark Generative Probabilistic Model for Weak Supervised Learning

Finding relevant and high-quality datasets to train machine learning mod...
research
08/12/2022

Online Discovery of Evolving Groups over Massive-Scale Trajectory Streams

The increasing pervasiveness of object tracking technologies leads to hu...
research
02/28/2019

Improving fraud prediction with incremental data balancing technique for massive data streams

The performance of classification algorithms with a massive and highly i...
research
10/15/2020

Aerodynamic Data Predictions Based on Multi-task Learning

The quality of datasets is one of the key factors that affect the accura...

Please sign up or login with your details

Forgot password? Click here to reset