Finding citations for PubMed: A large-scale comparison between five freely available bibliographic data sources

10/30/2021
by   Zhentao Liang, et al.
0

As an important biomedical database, PubMed provides users with free access to abstracts of its documents. However, citations between these documents need to be collected from external data sources. Although previous studies have investigated the coverage of various data sources, the quality of citations is underexplored. In response, this study compares the coverage and citation quality of five freely available data sources on 30 million PubMed documents, including OpenCitations Index of CrossRef open DOI-to-DOI citations (COCI), Dimensions, Microsoft Academic Graph (MAG), National Institutes of Health Open Citation Collection (NIH-OCC), and Semantic Scholar Open Research Corpus (S2ORC). Three gold standards and five metrics are introduced to evaluate the correctness and completeness of citations. Our results indicate that Dimensions is the most comprehensive data source that provides references for 62.4 PubMed documents, outperforming the official NIH-OCC dataset (56.7 of citation links in other data sources can also be found in Dimensions. The coverage of MAG, COCI, and S2ORC is 59.6 Regarding the citation quality, Dimensions and NIH-OCC achieve the best overall results. Almost all data sources have a precision higher than 90 recall is much lower. All databases have better performances on recent publications than earlier ones. Meanwhile, the gaps between different data sources have diminished for the documents published in recent years. This study provides evidence for researchers to choose suitable PubMed citation sources, which is also helpful for evaluating the citation quality of free bibliographic databases.

READ FULL TEXT
research
05/21/2020

Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic

We present a large-scale comparison of five multidisciplinary bibliograp...
research
08/27/2021

A map of Digital Humanities research across bibliographic data sources

Purpose. This study presents the results of an experiment we performed t...
research
03/16/2018

Evidence of Open Access of scientific publications in Google Scholar: a large-scale analysis

This article uses Google Scholar (GS) as a source of data to analyse Ope...
research
08/22/2018

Reproducible data citations for computational research

The general purpose of a scientific publication is the exchange and spre...
research
10/13/2021

Refcat: The Internet Archive Scholar Citation Graph

As part of its scholarly data efforts, the Internet Archive (IA) release...
research
10/01/2021

The case for the Humanities Citation Index (HuCI): a citation index by the humanities, for the humanities

Citation indexes are by now part of the research infrastructure in use b...

Please sign up or login with your details

Forgot password? Click here to reset