MetaEnhance: Metadata Quality Improvement for Electronic Theses and Dissertations of University Libraries

03/30/2023
by   Muntabir Hasan Choudhury, et al.
0

Metadata quality is crucial for digital objects to be discovered through digital library interfaces. However, due to various reasons, the metadata of digital objects often exhibits incomplete, inconsistent, and incorrect values. We investigate methods to automatically detect, correct, and canonicalize scholarly metadata, using seven key fields of electronic theses and dissertations (ETDs) as a case study. We propose MetaEnhance, a framework that utilizes state-of-the-art artificial intelligence methods to improve the quality of these fields. To evaluate MetaEnhance, we compiled a metadata quality evaluation benchmark containing 500 ETDs, by combining subsets sampled using multiple criteria. We tested MetaEnhance on this benchmark and found that the proposed methods achieved nearly perfect F1-scores in detecting errors and F1-scores in correcting errors ranging from 0.85 to 1.00 for five of seven fields.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/03/2017

Metadata in the BioSample Online Repository are Impaired by Numerous Anomalies

The metadata about scientific experiments are crucial for finding, repro...
research
08/17/2018

The Variable Quality of Metadata About Biological Samples Used in Biomedical Experiments

We present an analytical study of the quality of metadata about samples ...
research
05/15/2023

Representing provenance and track changes of cultural heritage metadata in RDF: a survey of existing approaches

The data within collections from all Digital Humanities fields must be t...
research
04/17/2018

Prioritizing and Scheduling Conferences for Metadata Harvesting in dblp

Maintaining literature databases and online bibliographies is a core res...
research
04/08/2021

It's All About The Cards: Sharing on Social Media Probably Encouraged HTML Metadata Growth

In a perfect world, all articles consistently contain sufficient metadat...
research
09/25/2020

A review of metadata fields associated with podcast RSS feeds

Podcasts are traditionally shared through RSS feeds. As well as pointing...
research
09/20/2022

Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics

Modern machine learning research relies on relatively few carefully cura...

Please sign up or login with your details

Forgot password? Click here to reset