Large scale biomedical texts classification: a kNN and an ESA-based approaches

06/09/2016
by   Khadim Dramé, et al.
0

With the large and increasing volume of textual data, automated methods for identifying significant topics to classify textual documents have received a growing interest. While many efforts have been made in this direction, it still remains a real challenge. Moreover, the issue is even more complex as full texts are not always freely available. Then, using only partial information to annotate these documents is promising but remains a very ambitious issue. MethodsWe propose two classification methods: a k-nearest neighbours (kNN)-based approach and an explicit semantic analysis (ESA)-based approach. Although the kNN-based approach is widely used in text classification, it needs to be improved to perform well in this specific classification problem which deals with partial information. Compared to existing kNN-based methods, our method uses classical Machine Learning (ML) algorithms for ranking the labels. Additional features are also investigated in order to improve the classifiers' performance. In addition, the combination of several learning algorithms with various techniques for fixing the number of relevant topics is performed. On the other hand, ESA seems promising for this classification task as it yielded interesting results in related issues, such as semantic relatedness computation between texts and text classification. Unlike existing works, which use ESA for enriching the bag-of-words approach with additional knowledge-based features, our ESA-based method builds a standalone classifier. Furthermore, we investigate if the results of this method could be useful as a complementary feature of our kNN-based approach.ResultsExperimental evaluations performed on large standard annotated datasets, provided by the BioASQ organizers, show that the kNN-based method with the Random Forest learning algorithm achieves good performances compared with the current state-of-the-art methods, reaching a competitive f-measure of 0.55 yielded reserved results.ConclusionsWe have proposed simple classification methods suitable to annotate textual documents using only partial information. They are therefore adequate for large multi-label classification and particularly in the biomedical domain. Thus, our work contributes to the extraction of relevant information from unstructured documents in order to facilitate their automated processing. Consequently, it could be used for various purposes, including document indexing, information retrieval, etc.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/13/2018

ML-Net: multi-label classification of biomedical texts with deep neural networks

Background: Multi-label text classification is one type of text classifi...
research
01/20/2018

Using Deep Learning For Title-Based Semantic Subject Indexing To Reach Competitive Performance to Full-Text

For (semi-)automated subject indexing systems in digital libraries, it i...
research
05/04/2023

Enhancing Pashto Text Classification using Language Processing Techniques for Single And Multi-Label Analysis

Text classification has become a crucial task in various fields, leading...
research
04/18/2017

Large-Scale Online Semantic Indexing of Biomedical Articles via an Ensemble of Multi-Label Classification Models

Background: In this paper we present the approaches and methods employed...
research
08/25/2023

Compressor-Based Classification for Atrial Fibrillation Detection

Atrial fibrillation (AF) is one of the most common arrhythmias with chal...
research
05/12/2021

Priberam at MESINESP Multi-label Classification of Medical Texts Task

Medical articles provide current state of the art treatments and diagnos...
research
03/02/2023

Document Provenance and Authentication through Authorship Classification

Style analysis, which is relatively a less explored topic, enables sever...

Please sign up or login with your details

Forgot password? Click here to reset