DeepAI AI Chat
Log In Sign Up

An Attention Ensemble Approach for Efficient Text Classification of Indian Languages

by   Atharva Kulkarni, et al.

The recent surge of complex attention-based deep learning architectures has led to extraordinary results in various downstream NLP tasks in the English language. However, such research for resource-constrained and morphologically rich Indian vernacular languages has been relatively limited. This paper proffers team SPPU_AKAH's solution for the TechDOfication 2020 subtask-1f: which focuses on the coarse-grained technical domain identification of short text documents in Marathi, a Devanagari script-based Indian language. Availing the large dataset at hand, a hybrid CNN-BiLSTM attention ensemble model is proposed that competently combines the intermediate sentence representations generated by the convolutional neural network and the bidirectional long short-term memory, leading to efficient text classification. Experimental results show that the proposed model outperforms various baseline machine learning and deep learning models in the given task, giving the best validation accuracy of 89.57% and f1-score of 0.8875. Furthermore, the solution resulted in the best system submission for this subtask, giving a test accuracy of 64.26% and f1-score of 0.6157, transcending the performances of other teams as well as the baseline system given by the organizers of the shared task.


page 1

page 2

page 3

page 4


TechTexC: Classification of Technical Texts using Convolution and Bidirectional Long Short Term Memory Network

This paper illustrates the details description of technical text classif...

Predicting Organizational Cybersecurity Risk: A Deep Learning Approach

Cyberattacks conducted by malicious hackers cause irreparable damage to ...

muBoost: An Effective Method for Solving Indic Multilingual Text Classification Problem

Text Classification is an integral part of many Natural Language Process...

Efficient Urdu Caption Generation using Attention based LSTMs

Recent advancements in deep learning has created a lot of opportunities ...

Joint Optimization of Tokenization and Downstream Model

Since traditional tokenizers are isolated from a downstream task and mod...

Cross-Lingual Task-Specific Representation Learning for Text Classification in Resource Poor Languages

Neural network models have shown promising results for text classificati...