EPIC: An Epidemics Corpus Of Over 20 Million Relevant Tweets

06/09/2020
by   Junhua Liu, et al.
0

Since the start of COVID-19, several relevant corpora from various sources are presented in the literature that contain millions of data points. While these corpora are valuable in supporting many analyses on this specific pandemic, researchers require additional benchmark corpora that contain other epidemics to facilitate cross-epidemic pattern recognition and trend analysis tasks. During our other efforts on COVID-19 related work, we discover very little disease related corpora in the literature that are sizable and rich enough to support such cross-epidemic analysis tasks. In this paper, we present EPIC, a large-scale epidemic corpus that contains 20 millions micro-blog posts, i.e., tweets crawled from Twitter, from year 2006 to 2020. EPIC contains a subset of 17.8 millions tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 3.5 millions tweets of six global epidemic outbreaks, including 2009 H1N1 Swine Flu, 2010 Haiti Cholera, 2012 Middle-East Respiratory Syndrome (MERS), 2013 West African Ebola, 2016 Yemen Cholera and 2018 Kivu Ebola. Furthermore, we explore and discuss the properties of the corpus with statistics of key terms and hashtags and trends analysis for each subset. Finally, we demonstrate the value and impact that EPIC could create through a discussion of multiple use cases of cross-epidemic research topics that attract growing interest in recent years. These use cases span multiple research areas, such as epidemiological modeling, pattern recognition, natural language understanding and economical modeling.

READ FULL TEXT

page 3

page 4

research
06/09/2020

EPIC30M: An Epidemics Corpus Of Over 30 Million Relevant Tweets

Since the start of COVID-19, several relevant corpora from various sourc...
research
07/10/2020

What Can We Learn From Almost a Decade of Food Tweets

We present the Latvian Twitter Eater Corpus - a set of tweets in the nar...
research
06/25/2020

TweetsCOV19 – A Knowledge Base of Semantically Annotated Tweets about the COVID-19 Pandemic

Publicly available social media archives facilitate research in the soci...
research
05/22/2020

GeoCoV19: A Dataset of Hundreds of Millions of Multilingual COVID-19 Tweets with Location Information

The past several years have witnessed a huge surge in the use of social ...
research
02/04/2020

Plague Dot Text: Text mining and annotation of outbreak reports of the Third Plague Pandemic (1894-1952)

The design of models that govern diseases in population is commonly buil...
research
10/10/2021

An Analysis of COVID-19 Knowledge Graph Construction and Applications

The construction and application of knowledge graphs have seen a rapid i...
research
07/10/2021

Computational Paremiology: Charting the temporal, ecological dynamics of proverb use in books, news articles, and tweets

Proverbs are an essential component of language and culture, and though ...

Please sign up or login with your details

Forgot password? Click here to reset