The Variable Quality of Metadata About Biological Samples Used in Biomedical Experiments

08/17/2018
by   Rafael S. Gonçalves, et al.
0

We present an analytical study of the quality of metadata about samples used in biomedical experiments. The metadata under analysis are stored in two well- known databases: BioSample---a repository managed by the National Center for Biotechnology Information (NCBI), and BioSamples---a repository managed by the European Bioinformatics Institute (EBI). We tested whether 11.4M sample metadata records in the two repositories are populated with values that fulfill the stated requirements for such values. Our study revealed multiple anomalies in the metadata. Most metadata field names and their values are not standardized or controlled. Even simple binary or numeric fields are often populated with inadequate values of different data types. By clustering metadata field names, we discovered there are often many distinct ways to represent the same aspect of a sample. Overall, the metadata we analyzed reveal that there is a lack of principled mechanisms to enforce and validate metadata requirements. The significant aberrancies that we found in the metadata are likely to impede search and secondary use of the associated datasets.

READ FULL TEXT
research
08/03/2017

Metadata in the BioSample Online Repository are Impaired by Numerous Anomalies

The metadata about scientific experiments are crucial for finding, repro...
research
03/19/2019

Aligning Biomedical Metadata with Ontologies Using Clustering and Embeddings

The metadata about scientific experiments published in online repositori...
research
03/21/2019

Using association rule mining and ontologies to generate metadata recommendations from multiple biomedical databases

Metadata-the machine-readable descriptions of the data-are increasingly ...
research
03/30/2023

MetaEnhance: Metadata Quality Improvement for Electronic Theses and Dissertations of University Libraries

Metadata quality is crucial for digital objects to be discovered through...
research
09/25/2020

A review of metadata fields associated with podcast RSS feeds

Podcasts are traditionally shared through RSS feeds. As well as pointing...
research
02/27/2020

Dataset Search In Biodiversity Research: Do Metadata In Data Repositories Reflect Scholarly Information Needs?

The increasing amount of research data provides the opportunity to link ...
research
11/08/2022

The French National 3D Data Repository for Humanities: Features, Feedback and Open Questions

We introduce the French National 3D Data Repository for Humanities desig...

Please sign up or login with your details

Forgot password? Click here to reset