Cleaning Noisy and Heterogeneous Metadata for Record Linking Across Scholarly Big Datasets

06/20/2019
by   Athar Sefid, et al.
0

Automatically extracted metadata from scholarly documents in PDF formats is usually noisy and heterogeneous, often containing incomplete fields and erroneous values. One common way of cleaning metadata is to use a bibliographic reference dataset. The challenge is to match records between corpora with high precision. The existing solution which is based on information retrieval and string similarity on titles works well only if the titles are cleaned. We introduce a system designed to match scholarly document entities with noisy metadata against a reference dataset. The blocking function uses the classic BM25 algorithm to find the matching candidates from the reference data that has been indexed by ElasticSearch. The core components use supervised methods which combine features extracted from all available metadata fields. The system also leverages available citation information to match entities. The combination of metadata and citation achieves high accuracy that significantly outperforms the baseline method on the same test dataset. We apply this system to match the database of CiteSeerX against Web of Science, PubMed, and DBLP. This method will be deployed in the CiteSeerX system to clean metadata and link records to other scholarly big datasets.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/08/2022

How to structure citations data and bibliographic metadata in the OpenCitations accepted format

The OpenCitations organization is working on ingesting citation data and...
research
05/26/2022

The way we cite: common metadata used across disciplines for defining bibliographic references

Current citation practices observed in articles are very noisy, confusin...
research
08/03/2017

Metadata in the BioSample Online Repository are Impaired by Numerous Anomalies

The metadata about scientific experiments are crucial for finding, repro...
research
10/27/2017

New Methods for Metadata Extraction from Scientific Literature

Within the past few decades we have witnessed digital revolution, which ...
research
11/26/2018

ParsRec: A Novel Meta-Learning Approach to Recommending Bibliographic Reference Parsers

Bibliographic reference parsers extract machine-readable metadata such a...
research
03/17/2019

Shining a light on Spotlight: Leveraging Apple's desktop search utility to recover deleted file metadata on macOS

Spotlight is a proprietary desktop search technology released by Apple i...
research
04/17/2018

Prioritizing and Scheduling Conferences for Metadata Harvesting in dblp

Maintaining literature databases and online bibliographies is a core res...

Please sign up or login with your details

Forgot password? Click here to reset