Old Content and Modern Tools - Searching Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910

11/09/2016
by   Kimmo Kettunen, et al.
0

Named Entity Recognition (NER), search, classification and tagging of names and name like frequent informational elements in texts, has become a standard information extraction procedure for textual data. NER has been applied to many types of texts and different types of entities: newspapers, fiction, historical records, persons, locations, chemical compounds, protein families, animals etc. In general a NER system's performance is genre and domain dependent and also used entity categories vary (Nadeau and Sekine, 2007). The most general set of named entities is usually some version of three partite categorization of locations, persons and organizations. In this paper we report first large scale trials and evaluation of NER with data out of a digitized Finnish historical newspaper collection Digi. Experiments, results and discussion of this research serve development of the Web collection of historical Finnish newspapers. Digi collection contains 1,960,921 pages of newspaper material from years 1771-1910 both in Finnish and Swedish. We use only material of Finnish documents in our evaluation. The OCRed newspaper collection has lots of OCR errors; its estimated word level correctness is about 70-75 Pääkkönen, 2016). Our principal NER tagger is a rule-based tagger of Finnish, FiNER, provided by the FIN-CLARIN consortium. We show also results of limited category semantic tagging with tools of the Semantic Computing Research Group (SeCo) of the Aalto University. Three other tools are also evaluated briefly. This research reports first published large scale results of NER in a historical Finnish OCRed newspaper collection. Results of the research supplement NER results of other languages with similar noisy data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/20/2023

A Benchmark of Nested Named Entity Recognition Approaches in Historical Structured Documents

Named Entity Recognition (NER) is a key step in the creation of structur...
research
12/30/2016

PAMPO: using pattern matching and pos-tagging for effective Named Entities recognition in Portuguese

This paper deals with the entity extraction task (named entity recogniti...
research
12/06/2018

Neural Word Search in Historical Manuscript Collections

We address the problem of segmenting and retrieving word images in colle...
research
05/31/2022

hmBERT: Historical Multilingual Language Models for Named Entity Recognition

Compared to standard Named Entity Recognition (NER), identifying persons...
research
09/23/2021

Named Entity Recognition and Classification on Historical Documents: A Survey

After decades of massive digitisation, an unprecedented amount of histor...
research
02/14/2016

Exploiting Lists of Names for Named Entity Identification of Financial Institutions from Unstructured Documents

There is a wealth of information about financial systems that is embedde...
research
01/30/2018

A Machine Learning Approach to Quantitative Prosopography

Prosopography is an investigation of the common characteristics of a gro...

Please sign up or login with your details

Forgot password? Click here to reset