Restoring and Mining the Records of the Joseon Dynasty via Neural Language Modeling and Machine Translation

04/13/2021
by   Kyeongpil Kang, et al.
0

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since most of the documents are not written in a modern language and part of the contents are damaged over time. As a result, restoring the damaged or unrecognizable parts as well as translating the records into modern languages are crucial tasks. In response, we present a multi-task learning approach to restore and translate historical documents based on a self-attention mechanism, specifically utilizing two Korean historical records, ones of the most voluminous historical records in the world. Experimental results show that our approach significantly improves the accuracy of the translation task than baselines without multi-task learning. In addition, we present an in-depth exploratory analysis on our translated results via topic modeling, uncovering several significant historical events.

READ FULL TEXT

page 2

page 9

research
05/20/2022

Translating Hanja historical documents to understandable Korean and English

The Annals of Joseon Dynasty (AJD) contain the daily records of the King...
research
06/28/2022

Placing (Historical) Facts on a Timeline: A Classification cum Coref Resolution Approach

A timeline provides one of the most effective ways to visualize the impo...
research
12/20/2017

Mining Events with Declassified Diplomatic Documents

Since 1973 the State Department has been using electronic records system...
research
10/25/2016

Improving historical spelling normalization with bi-directional LSTMs and multi-task learning

Natural-language processing of historical documents is complicated by th...
research
09/04/2020

Externalizing Transformations of Historical Documents: Opportunities for Provenance-Driven Visualization

Transcription, annotation, digitization and/or visualization are common ...
research
02/27/2019

Ranking in Genealogy: Search Results Fusion at Ancestry

Genealogy research is the study of family history using available resour...
research
07/01/2019

Modernizing Historical Documents: a User Study

Accessibility to historical documents is mostly limited to scholars. Thi...

Please sign up or login with your details

Forgot password? Click here to reset