Open Reproducible Publication Research

by   Diomidis Spinellis, et al.

Considerable scientific work involves locating, analyzing, systematizing, and synthesizing other publications. Its results end up in a paper's "background" section or in standalone articles, which include meta-analyses and systematic literature reviews. The required research is aided through the use of online scientific publication databases and search engines, such as Web of Science, Scopus, and Google Scholar. However, use of online databases suffers from a lack of repeatability and transparency, as well as from technical restrictions. Thankfully, open data, powerful personal computers, and open source software now make it possible to run sophisticated publication studies on the desktop in a self-contained environment that peers can readily reproduce. Here we report a Python software package and an associated command-line tool that can populate embedded relational databases with slices from the complete set of Crossref publication metadata, ORCID author records, and other open data sets, for in-depth processing through performant queries. We demonstrate the software's utility by analyzing the underlying dataset's contents, by visualizing the evolution of publications in diverse scientific fields and relationships among them, by outlining scientometric facts associated with COVID-19 research, and by replicating commonly-used bibliometric measures of productivity, impact, and disruption.


Do neutrons publish? A neutron publication survey 2005-2015

Publication in scientific journals is the main product of scientific res...

Generating large-scale network analyses of scientific landscapes in seconds using Dimensions on Google BigQuery

The growth of large, programatically accessible bibliometrics databases ...

Reproducible Research is more than Publishing Research Artefacts: A Systematic Analysis of Jupyter Notebooks from Research Articles

With the advent of Open Science, researchers have started to publish the...

Linking Mathematical Software in Web Archives

The Web is our primary source of all kinds of information today. This in...

pylustrator: Code generation for reproducible figures for publication

One major challenge in science is to make all results potentially reprod...

The 'Problematic Paper Screener' automatically selects suspect publications for post-publication (re)assessment

Post publication assessment remains necessary to check erroneous or frau...

Ribonucleic acid (RNA) virus and coronavirus in Google Dataset Search: their scope and epidemiological correlation

This paper presents an analysis of the publication of datasets collected...

Please sign up or login with your details

Forgot password? Click here to reset