VoynaSlov: A Data Set of Russian Social Media Activity during the 2022 Ukraine-Russia War

05/24/2022
by   Chan Young Park, et al.
0

In this report, we describe a new data set called VoynaSlov which contains 21M+ Russian-language social media activities (i.e. tweets, posts, comments) made by Russian media outlets and by the general public during the time of war between Ukraine and Russia. We scraped the data from two major platforms that are widely used in Russia: Twitter and VKontakte (VK), a Russian social media platform based in Saint Petersburg commonly referred to as "Russian Facebook". We provide descriptions of our data collection process and data statistics that compare state-affiliated and independent Russian media, and also the two platforms, VK and Twitter. The main differences that distinguish our data from previously released data related to the ongoing war are its focus on Russian media and consideration of state-affiliation as well as the inclusion of data from VK, which is more suitable than Twitter for understanding Russian public sentiment considering its wide use within Russia. We hope our data set can facilitate future research on information warfare and ultimately enable the reduction and prevention of disinformation and opinion manipulation campaigns. The data set is available at https://github.com/chan0park/VoynaSlov and will be regularly updated as we continuously collect more data.

READ FULL TEXT

page 6

page 11

research
01/18/2021

Capitol (Pat)riots: A comparative study of Twitter and Parler

On 6 January 2021, a mob of right-wing conservatives stormed the USA Cap...
research
03/14/2022

Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War Between Ukraine and Russia

On February 24, 2022, Russia invaded Ukraine. In the days that followed,...
research
05/05/2020

Reliable and Efficient Long-Term Twitter Monitoring

Social media data is now widely used by many academic researchers. Howev...
research
01/23/2020

The Pushshift Reddit Dataset

Social media data has become crucial to the advancement of scientific un...
research
07/13/2023

Electoral Agitation Data Set: The Use Case of the Polish Election

The popularity of social media makes politicians use it for political ad...
research
08/07/2017

FixMyStreet Brussels: Socio-Demographic Inequality in Crowdsourced Civic Participation

FixMyStreet (FMS) is a web-based civic participation platform that allow...
research
04/13/2023

Vax-Culture: A Dataset for Studying Vaccine Discourse on Twitter

Vaccine hesitancy continues to be a main challenge for public health off...

Please sign up or login with your details

Forgot password? Click here to reset