Wikipedia Citations: A comprehensive dataset of citations with identifiers extracted from English Wikipedia

07/14/2020
by   Harshdeep Singh, et al.
0

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close this gap, we release Wikipedia Citations, a comprehensive dataset of citations extracted from Wikipedia. A total of 29.3M citations were extracted from 6.1M English Wikipedia articles as of May 2020, and classified as being to books, journal articles or Web contents. We were thus able to extract 4.0M citations to scholarly publications with known identifiers – including DOI, PMC, PMID, and ISBN – and further equip an extra 261K citations with DOIs from Crossref. As a result, we find that 6.7 one journal article with an associated DOI, and that Wikipedia cites just 2 all articles with a DOI currently indexed in the Web of Science. We release our code to allow the community to extend upon our work and update the dataset in the future.

READ FULL TEXT
research
10/26/2021

A Map of Science in Wikipedia

In recent decades, the rapid growth of Internet adoption is offering opp...
research
10/28/2022

Polarization and reliability of news sources in Wikipedia

Wikipedia is the largest online encyclopedia: its open contribution poli...
research
01/23/2020

Quantifying Engagement with Citations on Wikipedia

Wikipedia, the free online encyclopedia that anyone can edit, is one of ...
research
02/28/2019

Citation Needed: A Taxonomy and Algorithmic Assessment of Wikipedia's Verifiability

Wikipedia is playing an increasingly central role on the web,and the pol...
research
05/23/2023

Wikipedia and open access

Wikipedia is a well-known platform for disseminating knowledge, and scie...
research
03/30/2017

Finding News Citations for Wikipedia

An important editing policy in Wikipedia is to provide citations for add...
research
10/13/2021

Refcat: The Internet Archive Scholar Citation Graph

As part of its scholarly data efforts, the Internet Archive (IA) release...

Please sign up or login with your details

Forgot password? Click here to reset