HistRED: A Historical Document-Level Relation Extraction Dataset

07/10/2023
by   Soyoung Yang, et al.
0

Despite the extensive applications of relation extraction (RE) tasks in various domains, little has been explored in the historical context, which contains promising data across hundreds and thousands of years. To promote the historical RE research, we present HistRED constructed from Yeonhaengnok. Yeonhaengnok is a collection of records originally written in Hanja, the classical Chinese writing, which has later been translated into Korean. HistRED provides bilingual annotations such that RE can be performed on Korean and Hanja texts. In addition, HistRED supports various self-contained subtexts with different lengths, from a sentence level to a document level, supporting diverse context settings for researchers to evaluate the robustness of their RE models. To demonstrate the usefulness of our dataset, we propose a bilingual RE model that leverages both Korean and Hanja contexts to predict relations between entities. Our model outperforms monolingual baselines on HistRED, showing that employing multiple language contexts supplements the RE predictions. The dataset is publicly available at: https://huggingface.co/datasets/Soyoung/HistRED under CC BY-NC-ND 4.0 license.

READ FULL TEXT

page 1

page 6

page 8

page 13

research
08/21/2021

A Hierarchical Entity Graph Convolutional Network for Relation Extraction across Documents

Distantly supervised datasets for relation extraction mostly focus on se...
research
10/19/2022

CEntRE: A paragraph-level Chinese dataset for Relation Extraction among Enterprises

Enterprise relation extraction aims to detect pairs of enterprise entiti...
research
02/20/2021

Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation Extraction

Entities, as the essential elements in relation extraction tasks, exhibi...
research
09/29/2022

TERMinator: A system for scientific texts processing

This paper is devoted to the extraction of entities and semantic relatio...
research
05/25/2022

Revisiting DocRED – Addressing the Overlooked False Negative Problem in Relation Extraction

The DocRED dataset is one of the most popular and widely used benchmarks...
research
06/20/2023

Did the Models Understand Documents? Benchmarking Models for Language Understanding in Document-Level Relation Extraction

Document-level relation extraction (DocRE) attracts more research intere...
research
03/08/2019

ICDAR 2019 Historical Document Reading Challenge on Large Structured Chinese Family Records

We propose a Historical Document Reading Challenge on Large Chinese Stru...

Please sign up or login with your details

Forgot password? Click here to reset