Structured dataset documentation: a datasheet for CheXpert

05/07/2021
by   Christian Garbin, et al.
43

Billions of X-ray images are taken worldwide each year. Machine learning, and deep learning in particular, has shown potential to help radiologists triage and diagnose images. However, deep learning requires large datasets with reliable labels. The CheXpert dataset was created with the participation of board-certified radiologists, resulting in the strong ground truth needed to train deep learning networks. Following the structured format of Datasheets for Datasets, this paper expands on the original CheXpert paper and other sources to show the critical role played by radiologists in the creation of reliable labels and to describe the different aspects of the dataset composition in detail. Such structured documentation intends to increase the awareness in the machine learning and medical communities of the strengths, applications, and evolution of CheXpert, thereby advancing the field of medical image analysis. Another objective of this paper is to put forward this dataset datasheet as an example to the community of how to create detailed and structured descriptions of datasets. We believe that clearly documenting the creation process, the contents, and applications of datasets accelerates the creation of useful and reliable models.

READ FULL TEXT

page 7

page 8

page 10

research
12/05/2019

Deep learning with noisy labels: exploring techniques and remedies in medical image analysis

Supervised training of deep learning models requires large labeled datas...
research
06/07/2017

Synthesizing Filamentary Structured Images with GANs

This paper aims at synthesizing filamentary structured images such as re...
research
02/15/2019

Going Deep in Medical Image Analysis: Concepts, Methods, Challenges and Future Directions

Medical Image Analysis is currently experiencing a paradigm shift due to...
research
07/11/2017

Creatism: A deep-learning photographer capable of creating professional work

Machine-learning excels in many areas with well-defined goals. However, ...
research
04/22/2021

Colonoscopy Polyp Detection and Classification: Dataset Creation and Comparative Evaluations

Colorectal cancer (CRC) is one of the most common types of cancer with a...
research
12/22/2021

Community Detection in Medical Image Datasets: Using Wavelets and Spectral Methods

Medical image datasets can have large number of images representing pati...
research
09/27/2021

Machine Learning based Medical Image Deepfake Detection: A Comparative Study

Deep generative networks in recent years have reinforced the need for ca...

Please sign up or login with your details

Forgot password? Click here to reset