Packaging code for reproducible research in the public sector

by   Federico Botta, et al.

The effective and ethical use of data to inform decision-making offers huge value to the public sector, especially when delivered by transparent, reproducible, and robust data processing workflows. One way that governments are unlocking this value is through making their data publicly available, allowing more people and organisations to derive insights. However, open data is not enough in many cases: publicly available datasets need to be accessible in an analysis-ready form from popular data science tools, such as R and Python, for them to realise their full potential. This paper explores ways to maximise the impact of open data with reference to a case study of packaging code to facilitate reproducible analysis. We present the jtstats project, which consists of R and Python packages for importing, processing, and visualising large and complex datasets representing journey times, for many modes and purposes at multiple geographic levels, released by the UK Department of Transport. jtstats shows how domain specific packages can enable reproducible research within the public sector and beyond, saving duplicated effort and reducing the risks of errors from repeated analyses. We hope that the jtstats project inspires others, particularly those in the public sector, to add value to their data sets by making them more accessible.


FluidDyn: a Python open-source framework for research and teaching in fluid dynamics

FluidDyn is a project to foster open-science and open-source in the flui...

A Framework For Sharing Publicly Available Data To Inform The COVID-19 Outbreak in Africa: A South African Case Study

The coronavirus disease (COVID-19), caused by the SARS-CoV-2 virus, beca...

One DSL to Rule Them All: IDE-Assisted Code Generation for Agile Data Analysis

Data analysis is at the core of scientific studies, a prominent task tha...

A Reflection on the Structure and Process of the Web of Data

The Web community has introduced a set of standards and technologies for...

Blended Integrated Open Data: dados abertos públicos integrados

While several public institutions provide its data openly, the effort re...

Towards Data-Driven Precision Agriculture using Open Data and Open Source Software

Information and communications technology (ICT) within the agricultural ...

Playing catch-up in building an open research commons

On August 2, 2021 a group of concerned scientists and US funding agency ...

Please sign up or login with your details

Forgot password? Click here to reset