Librarian-in-the-Loop: A Natural Language Processing Paradigm for Detecting Informal Mentions of Research Data in Academic Literature

03/10/2022
by   Lizhou Fan, et al.
0

Data citations provide a foundation for studying research data impact. Collecting and managing data citations is a new frontier in archival science and scholarly communication. However, the discovery and curation of research data citations is labor intensive. Data citations that reference unique identifiers (i.e. DOIs) are readily findable; however, informal mentions made to research data are more challenging to infer. We propose a natural language processing (NLP) paradigm to support the human task of identifying informal mentions made to research datasets. The work of discovering informal data mentions is currently performed by librarians and their staff in the Inter-university Consortium for Political and Social Research (ICPSR), a large social science data archive that maintains a large bibliography of data-related literature. The NLP model is bootstrapped from data citations actively collected by librarians at ICPSR. The model combines pattern matching with multiple iterations of human annotations to learn additional rules for detecting informal data mentions. These examples are then used to train an NLP pipeline. The librarian-in-the-loop paradigm is centered in the data work performed by ICPSR librarians, supporting broader efforts to build a more comprehensive bibliography of data-related literature that reflects the scholarly communities of research data users.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/23/2022

A Natural Language Processing Pipeline for Detecting Informal Data References in Academic Literature

Discovering authoritative links between publications and the datasets th...
research
03/06/2021

Putting Humans in the Natural Language Processing Loop: A Survey

How can we design Natural Language Processing (NLP) systems that learn f...
research
01/31/2023

Breaking Out of the Ivory Tower: A Large-scale Analysis of Patent Citations to HCI Research

What is the impact of human-computer interaction research on industry? W...
research
04/16/2021

Translational NLP: A New Paradigm and General Principles for Natural Language Processing Research

Natural language processing (NLP) research combines the study of univers...
research
07/21/2023

Who should I Collaborate with? A Comparative Study of Academia and Industry Research Collaboration in NLP

The goal of our research was to investigate the effects of collaboration...
research
01/08/2023

The State of Human-centered NLP Technology for Fact-checking

Misinformation threatens modern society by promoting distrust in science...
research
06/20/2017

Making visible the invisible through the analysis of acknowledgements in the humanities

Purpose: Science is subject to a normative structure that includes how t...

Please sign up or login with your details

Forgot password? Click here to reset