Machine Learning Based Detection of Clickbait Posts in Social Media

10/05/2017
by   Xinyue Cao, et al.
0

Clickbait (headlines) make use of misleading titles that hide critical information from or exaggerate the content on the landing target pages to entice clicks. As clickbaits often use eye-catching wording to attract viewers, target contents are often of low quality. Clickbaits are especially widespread on social media such as Twitter, adversely impacting user experience by causing immense dissatisfaction. Hence, it has become increasingly important to put forward a widely applicable approach to identify and detect clickbaits. In this paper, we make use of a dataset from the clickbait challenge 2017 (clickbait-challenge.com) comprising of over 21,000 headlines/titles, each of which is annotated by at least five judgments from crowdsourcing on how clickbait it is. We attempt to build an effective computational clickbait detection model on this dataset. We first considered a total of 331 features, filtered out many features to avoid overfitting and improve the running time of learning, and eventually selected the 60 most important features for our final model. Using these features, Random Forest Regression achieved the following results: MSE=0.035 MSE, Accuracy=0.82, and F1-sore=0.61 on the clickbait class.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/29/2019

Towards Automatic Bot Detection in Twitter for Health-related Tasks

With the increasing use of social media data for health-related research...
research
10/08/2017

Clickbait detection using word embeddings

Clickbait is a pejorative term describing web content that is aimed at g...
research
10/01/2017

Identifying Clickbait Posts on Social Media with an Ensemble of Linear Models

The purpose of a clickbait is to make a link so appealing that people cl...
research
02/08/2017

Social media mining for identification and exploration of health-related information from pregnant women

Widespread use of social media has led to the generation of substantial ...
research
06/10/2020

Heterogeneous Graph Attention Networks for Early Detection of Rumors on Twitter

With the rapid development of mobile Internet technology and the widespr...
research
07/07/2023

Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning

Human-annotated data plays a critical role in the fairness of AI systems...
research
04/14/2021

Influenza Surveillance using Search Engine, SNS, On-line Shopping, Q A Service and Past Flu Patients

Influenza, an infectious disease, causes many deaths worldwide. Predicti...

Please sign up or login with your details

Forgot password? Click here to reset