Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings

03/04/2019
by   Gerhard Wohlgenannt, et al.
0

In this research, we manually create high-quality datasets in the digital humanities domain for the evaluation of language models, specifically word embedding models. The first step comprises the creation of unigram and n-gram datasets for two fantasy novel book series for two task types each, analogy and doesn't-match. This is followed by the training of models on the two book series with various popular word embedding model types such as word2vec, GloVe, fastText, or LexVec. Finally, we evaluate the suitability of word embedding models for such specific relation extraction tasks in a situation of comparably small corpus sizes. In the evaluations, we also investigate and analyze particular aspects such as the impact of corpus term frequencies and task difficulty on accuracy. The datasets, and the underlying system and word embedding models are available on github and can be easily extended with new datasets and tasks, be used to reproduce the presented results, or be transferred to other domains.

READ FULL TEXT
research
03/07/2019

Creation and Evaluation of Datasets for Distributional Semantics Tasks in the Digital Humanities Domain

Word embeddings are already well studied in the general domain, usually ...
research
07/20/2015

How to Generate a Good Word Embedding?

We analyze three critical components of word embedding training: the mod...
research
03/04/2019

Russian Language Datasets in the Digitial Humanities Domain and Their Evaluation with Word Embeddings

In this paper, we present Russian language datasets in the digital human...
research
01/08/2019

Deconstructing Word Embeddings

A review of Word Embedding Models through a deconstructive approach reve...
research
10/04/2016

Chinese Event Extraction Using DeepNeural Network with Word Embedding

A lot of prior work on event extraction has exploited a variety of featu...
research
07/20/2020

The Geometry of Information Cocoon: Analyzing the Cultural Space with Word Embedding Models

Accompanied by the rapid development of digital media, the threat of inf...
research
05/10/2015

Improved Relation Extraction with Feature-Rich Compositional Embedding Models

Compositional embedding models build a representation (or embedding) for...

Please sign up or login with your details

Forgot password? Click here to reset