Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

10/29/2019
by   Neil Deshmukh, et al.
0

Although machine learning has become a powerful tool to augment doctors in clinical analysis, the immense amount of labeled data that is necessary to train supervised learning approaches burdens each development task as time and resource intensive. The vast majority of dense clinical information is stored in written reports, detailing pertinent patient information. The challenge with utilizing natural language data for standard model development is due to the complex nature of the modality. In this research, a model pipeline was developed to utilize an unsupervised approach to train an encoder-language model, a recurrent network, to generate document encodings; which then can be used as features passed into a decoder-classifier model that requires magnitudes less labeled data than previous approaches to differentiate between fine-grained disease classes accurately. The language model was trained on unlabeled radiology reports from the Massachusetts General Hospital Radiology Department (n=218,159) and terminated with a loss of 1.62. The classification models were trained on three labeled datasets of head CT studies of reported patients, presenting large vessel occlusion (n=1403), acute ischemic strokes (n=331), and intracranial hemorrhage (n=4350), to identify a variety of different findings directly from the radiology report data; resulting in AUCs of 0.98, 0.95, and 0.99, respectively, for the large vessel occlusion, acute ischemic stroke, and intracranial hemorrhage datasets. The output encodings are able to be used in conjunction with imaging data, to create models that can process a multitude of different modalities. The ability to automatically extract relevant features from textual data allows for faster model development and integration of textual modality, overall, allowing clinical reports to become a more viable input for more encompassing and accurate deep learning models.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/03/2018

Clinical Document Classification Using Labeled and Unlabeled Data Across Hospitals

Reviewing radiology reports in emergency departments is an essential but...
research
07/21/2021

An artificial intelligence natural language processing pipeline for information extraction in neuroradiology

The use of electronic health records in medical research is difficult be...
research
07/24/2021

A Real Use Case of Semi-Supervised Learning for Mammogram Classification in a Local Clinic of Costa Rica

The implementation of deep learning based computer aided diagnosis syste...
research
10/01/2018

Efficient and Accurate Abnormality Mining from Radiology Reports with Customized False Positive Reduction

Obtaining datasets labeled to facilitate model development is a challeng...
research
07/06/2020

Labeling of Multilingual Breast MRI Reports

Medical reports are an essential medium in recording a patient's conditi...
research
02/28/2021

CREATe: Clinical Report Extraction and Annotation Technology

Clinical case reports are written descriptions of the unique aspects of ...
research
12/01/2021

MOMO – Deep Learning-driven classification of external DICOM studies for PACS archivation

Patients regularly continue assessment or treatment in other facilities ...

Please sign up or login with your details

Forgot password? Click here to reset