Harmonise and integrate heterogeneous areal data with the R package arealDB

09/14/2019
by   Steffen Ehrmann, et al.
0

Areal data is a common data type to store information such as biodiversity inventories, socio-economic censuses or cadastral surveys. Many research questions require that areal data are integrated from multiple heterogeneous sources. Inconsistent concepts, terms, definitions, or messy tables makes data-wrangling an often tedious and error-prone process. A dedicated tool that assists in organising areal data still is lacking. Here, we introduce the R package arealDB that helps to harmonise and integrate heterogeneous areal data and associated geometries into a consistent database. The package is used to collect metadata, harmonise language and variable names, reshape messy into tidy data and integrate them in a standard data format. arealDB solves the specific problem of integrating disparate regional data sources on a given target variable, which may be published in different languages, with a different table arrangement or provided in various data formats. We guide the user step by step through the individual functions needed to integrate two such datasets using the example of the harvested area of soybean in Brazil and the USA. A database that has been built with arealDB is "tidy", and can thus be accessed easily with powerful and widespread tools such as the R meta-package tidyverse. Moreover, it is accompanied by provenance documentation that traces the full process of creation for each data point in the database. By offering easy-to-use tools for integrating areal data, arealDB promises substantial time-savings to database collation efforts, as well as quality-improvements to downstream scientific, monitoring, and management applications.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/24/2018

On-Demand Big Data Integration: A Hybrid ETL Approach for Reproducible Scientific Research

Scientific research requires access, analysis, and sharing of data that ...
research
09/16/2018

Semantic Interoperability Middleware Architecture for Heterogeneous Environmental Data Sources

Data heterogeneity hampers the effort to integrate and infer knowledge f...
research
12/21/2018

GaussianProcesses.jl: A Nonparametric Bayes package for the Julia Language

Gaussian processes are a class of flexible nonparametric Bayesian tools ...
research
07/23/2021

ArchaeoDAL: A Data Lake for Archaeological Data Management and Analytics

With new emerging technologies, such as satellites and drones, archaeolo...
research
01/20/2022

Use of Simulation Models for the Development of a Statistical Production Framework for Mobile Network Data with the simutils Package

We propose to use agent-based simulation models for the development of s...
research
02/08/2018

Praaline: Integrating Tools for Speech Corpus Research

This paper presents Praaline, an open-source software system for managin...

Please sign up or login with your details

Forgot password? Click here to reset