TNT-KID: Transformer-based Neural Tagger for Keyword Identification

03/20/2020
by   Matej Martinc, et al.
0

With growing amounts of available textual data, development of algorithms capable of automatic analysis, categorization and summarization of these data has become a necessity. In this research we present a novel algorithm for keyword identification, i.e., an extraction of one or multi-word phrases representing key aspects of a given document, called Transformer-based Neural Tagger for Keyword IDentification (TNT-KID). By adapting the transformer architecture for a specific task at hand and leveraging language model pretraining on a small domain specific corpus, the model is capable of overcoming deficiencies of both supervised and unsupervised state-of-the-art approaches to keyword extraction by offering competitive and robust performance on a variety of different datasets while requiring only a fraction of manually labeled data required by the best performing systems. This study also offers thorough error analysis with valuable insights into inner workings of the model and an ablation study measuring the influence of specific components of the keyword identification workflow on the overall performance.

READ FULL TEXT

page 1

page 16

page 17

research
04/11/2017

Automatic Keyword Extraction for Text Summarization: A Survey

In recent times, data is growing rapidly in every domain such as news, s...
research
01/31/2021

Extending Neural Keyword Extraction with TF-IDF tagset matching

Keyword extraction is the task of identifying words (or multi-word expre...
research
09/28/2022

Keyword Extraction from Short Texts with a Text-To-Text Transfer Transformer

The paper explores the relevance of the Text-To-Text Transfer Transforme...
research
11/11/2022

Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting

In this paper, we present a novel approach to adapt a sequence-to-sequen...
research
07/14/2018

Generating Synthetic Data for Neural Keyword-to-Question Models

Search typically relies on keyword queries, but these are often semantic...
research
05/19/2020

GLEAKE: Global and Local Embedding Automatic Keyphrase Extraction

Automated methods for granular categorization of large corpora of text d...
research
09/11/2023

Unsupervised Bias Detection in College Student Newspapers

This paper presents a pipeline with minimal human influence for scraping...

Please sign up or login with your details

Forgot password? Click here to reset