Document Classification for COVID-19 Literature

06/15/2020
by   Bernal Jimenez Gutierrez, et al.
0

The global pandemic has made it more important than ever to quickly and accurately retrieve relevant scientific literature for effective consumption by researchers in a wide range of fields. We provide an analysis of several multi-label document classification models on the LitCovid dataset, a growing collection of 8,000 research papers regarding the novel 2019 coronavirus. We find that pre-trained language models fine-tuned on this dataset outperform all other baselines and that the BioBERT and novel Longformer models surpass all others with almost equivalent micro-F1 and accuracy scores of around 81 69 these models as essential features of any system prepared to deal with an urgent situation like the current health crisis. Finally, we explore 50 errors made by the best performing models on LitCovid documents and find that they often (1) correlate certain labels too closely together and (2) fail to focus on discriminative sections of the articles; both of which are important issues to address in future work. Both data and code are available on GitHub.

READ FULL TEXT
research
04/08/2022

Towards Understanding Large-Scale Discourse Structures in Pre-Trained and Fine-Tuned Language Models

With a growing number of BERTology work analyzing different components o...
research
06/07/2023

Good Data, Large Data, or No Data? Comparing Three Approaches in Developing Research Aspect Classifiers for Biomedical Papers

The rapid growth of scientific publications, particularly during the COV...
research
08/15/2020

Label-Wise Document Pre-Training for Multi-Label Text Classification

A major challenge of multi-label text classification (MLTC) is to stimul...
research
04/19/2022

LitMC-BERT: transformer-based multi-label classification of biomedical literature with an application on COVID-19 literature curation

The rapid growth of biomedical literature poses a significant challenge ...
research
11/05/2022

Hierarchical Multi-Label Classification of Scientific Documents

Automatic topic classification has been studied extensively to assist ma...
research
03/24/2021

CSFCube – A Test Collection of Computer Science Research Articles for Faceted Query by Example

Query by Example is a well-known information retrieval task in which a d...
research
04/12/2021

WHOSe Heritage: Classification of UNESCO World Heritage "Outstanding Universal Value" Documents with Smoothed Labels

The UNESCO World Heritage List (WHL) is to identify the exceptionally va...

Please sign up or login with your details

Forgot password? Click here to reset