Corpus Conversion Service: A machine learning platform to ingest documents at scale [Poster abstract]

05/15/2018
by   Peter W J Staar, et al.
0

Over the past few decades, the amount of scientific articles and technical literature has increased exponentially in size. Consequently, there is a great need for systems that can ingest these documents at scale and make their content discoverable. Unfortunately, both the format of these documents (e.g. the PDF format or bitmap images) as well as the presentation of the data (e.g. complex tables) make the extraction of qualitative and quantitive data extremely challenging. We present a platform to ingest documents at scale which is powered by Machine Learning techniques and allows the user to train custom models on document collections. We show precision/recall results greater than 97 evidence for each of the microservices constituting the platform.

READ FULL TEXT

page 1

page 2

page 3

research
05/24/2018

Corpus Conversion Service: A Machine Learning Platform to Ingest Documents at Scale

Over the past few decades, the amount of scientific articles and technic...
research
08/30/2023

Large-scale data extraction from the UNOS organ donor documents

The scope of our study is all UNOS data of the USA organ donors since 20...
research
06/01/2022

Delivering Document Conversion as a Cloud Service with High Throughput and Responsiveness

Document understanding is a key business process in the data-driven econ...
research
07/11/2017

Leipzig Corpus Miner - A Text Mining Infrastructure for Qualitative Data Analysis

This paper presents the "Leipzig Corpus Miner", a technical infrastructu...
research
11/28/2021

CHARTER: heatmap-based multi-type chart data extraction

The digital conversion of information stored in documents is a great sou...
research
09/19/2023

Semi-automatic staging area for high-quality structured data extraction from scientific literature

In this study, we propose a staging area for ingesting new superconducto...
research
03/13/2021

Lightweight Selective Disclosure for Verifiable Documents on Blockchain

To achieve lightweight selective disclosure for protecting privacy of do...

Please sign up or login with your details

Forgot password? Click here to reset