SuMe: A Dataset Towards Summarizing Biomedical Mechanisms

05/10/2022
by   Mohaddeseh Bastan, et al.
0

Can language models read biomedical texts and explain the biomedical mechanisms discussed? In this work we introduce a biomedical mechanism summarization task. Biomedical studies often investigate the mechanisms behind how one entity (e.g., a protein or a chemical) affects another in a biological context. The abstracts of these publications often include a focused set of sentences that present relevant supporting statements regarding such relationships, associated experimental evidence, and a concluding sentence that summarizes the mechanism underlying the relationship. We leverage this structure and create a summarization task, where the input is a collection of sentences and the main entities in an abstract, and the output includes the relationship and a sentence that summarizes the mechanism. Using a small amount of manually labeled mechanism sentences, we train a mechanism sentence classifier to filter a large biomedical abstract collection and create a summarization dataset with 22k instances. We also introduce conclusion sentence generation as a pretraining task with 611k instances. We benchmark the performance of large bio-domain language models. We find that while the pretraining task help improves performance, the best model produces acceptable mechanism outputs in only 32 significant challenges in biomedical language understanding and summarization.

READ FULL TEXT

page 6

page 7

research
10/26/2022

BioNLI: Generating a Biomedical NLI Dataset Using Lexico-semantic Constraints for Adversarial Examples

Natural language inference (NLI) is critical for complex decision-making...
research
03/07/2019

Small-world networks for summarization of biomedical articles

In recent years, many methods have been developed to identify important ...
research
05/15/2023

Comparing Variation in Tokenizer Outputs Using a Series of Problematic and Challenging Biomedical Sentences

Background Objective: Biomedical text data are increasingly availabl...
research
08/06/2019

Clustering of Deep Contextualized Representations for Summarization of Biomedical Texts

In recent years, summarizers that incorporate domain knowledge into the ...
research
08/06/2019

Text Summarization in the Biomedical Domain

This chapter gives an overview of recent advances in the field of biomed...
research
05/10/2016

Different approaches for identifying important concepts in probabilistic biomedical text summarization

Automatic text summarization tools help users in biomedical domain to ac...
research
10/17/2017

PubMed 200k RCT: a Dataset for Sequential Sentence Classification in Medical Abstracts

We present PubMed 200k RCT, a new dataset based on PubMed for sequential...

Please sign up or login with your details

Forgot password? Click here to reset