SIMARA: a database for key-value information extraction from full pages

04/26/2023
by   Solène Tarride, et al.
0

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwritten documents that contain metadata describing older archives. They are stored in the National Archives of France and are used by archivists to identify and find archival documents. Each document is annotated at page-level, and contains seven fields to retrieve. The localization of each field is not available in such a way that this dataset encourages research on segmentation-free systems for information extraction. We propose a model based on the Transformer architecture trained for end-to-end information extraction and provide three sets for training, validation and testing, to ensure fair comparison with future works. The database is freely accessible at https://zenodo.org/record/7868059.

READ FULL TEXT

page 5

page 10

page 11

research
09/22/2020

Whole page recognition of historical handwriting

Historical handwritten documents guard an important part of human knowle...
research
04/26/2023

Key-value information extraction from full handwritten pages

We propose a Transformer-based approach for information extraction from ...
research
06/23/2021

ScanBank: A Benchmark Dataset for Figure Extraction from Scanned Electronic Theses and Dissertations

We focus on electronic theses and dissertations (ETDs), aiming to improv...
research
03/08/2018

Towards Knowledge Discovery from the Vatican Secret Archives. In Codice Ratio -- Episode 1: Machine Transcription of the Manuscripts

In Codice Ratio is a research project to study tools and techniques for ...
research
03/10/2021

DeepCPCFG: Deep Learning and Context Free Grammars for End-to-End Information Extraction

We combine deep learning and Conditional Probabilistic Context Free Gram...
research
12/01/2020

HORAE: an annotated dataset of books of hours

We introduce in this paper a new dataset of annotated pages from books o...
research
10/18/2017

MEDOC: a Python wrapper to load MEDLINE into a local MySQL database

Since the MEDLINE database was released, the number of documents indexed...

Please sign up or login with your details

Forgot password? Click here to reset