A frame semantic overview of NLP-based information extraction for cancer-related EHR notes

04/02/2019
by   Surabhi Datta, et al.
0

Objective: There is a lot of information about cancer in Electronic Health Record (EHR) notes that can be useful for biomedical research provided natural language processing (NLP) methods are available to extract and structure this information. In this paper, we present a scoping review of existing clinical NLP literature for cancer. Methods: We identified studies describing an NLP method to extract specific cancer-related information from EHR sources from PubMed, Google Scholar, ACL Anthology, and existing reviews. Two exclusion criteria were used in this study. We excluded articles where the extraction techniques used were too broad to be represented as frames and also where very low-level extraction methods were used. 79 articles were included in the final review. We organized this information according to frame semantic principles to help identify common areas of overlap and potential gaps. Results: Frames were created from the reviewed articles pertaining to cancer information such as cancer diagnosis, tumor description, cancer procedure, breast cancer diagnosis, prostate cancer diagnosis and pain in prostate cancer patients. These frames included both a definition as well as specific frame elements (i.e. extractable attributes). We found that cancer diagnosis was the most common frame among the reviewed papers (36 out of 79), with recent work focusing on extracting information related to treatment and breast cancer diagnosis. Conclusion: The list of common frames described in this paper identifies important cancer-related information extracted by existing NLP techniques and serves as a useful resource for future researchers requiring cancer information extracted from EHR notes. We also argue, due to the heavy duplication of cancer NLP systems, that a general purpose resource of annotated cancer frames and corresponding NLP tools would be valuable.

READ FULL TEXT
research
06/02/2023

Publicly available datasets of breast histopathology H E whole-slide images: A systematic review

Advancements in digital pathology and computing resources have made a si...
research
12/06/2017

An innovative solution for breast cancer textual big data analysis

The digitalization of stored information in hospitals now allows for the...
research
11/21/2022

Unsupervised extraction, labelling and clustering of segments from clinical notes

This work is motivated by the scarcity of tools for accurate, unsupervis...
research
08/07/2023

Extracting detailed oncologic history and treatment plan from medical oncology notes with large language models

Both medical care and observational studies in oncology require a thorou...
research
11/18/2019

Drug Repurposing for Cancer: An NLP Approach to Identify Low-Cost Therapies

More than 200 generic drugs approved by the U.S. Food and Drug Administr...
research
09/30/2020

Extracting Concepts for Precision Oncology from the Biomedical Literature

This paper describes an initial dataset and automatic natural language p...

Please sign up or login with your details

Forgot password? Click here to reset