Synthetic vs. Real Reference Strings for Citation Parsing, and the Importance of Re-training and Out-Of-Sample Data for Meaningful Evaluations: Experiments with GROBID, GIANT a

04/22/2020
by   Mark Grennan, et al.
2

Citation parsing, particularly with deep neural networks, suffers from a lack of training data as available datasets typically contain only a few thousand training instances. Manually labelling citation strings is very time-consuming, hence synthetically created training data could be a solution. However, as of now, it is unknown if synthetically created reference-strings are suitable to train machine learning algorithms for citation parsing. To find out, we train Grobid, which uses Conditional Random Fields, with a) human-labelled reference strings from 'real' bibliographies and b) synthetically created reference strings from the GIANT dataset. We find that both synthetic and organic reference strings are equally suited for training Grobid (F1 = 0.74). We additionally find that retraining Grobid has a notable impact on its performance, for both synthetic and real data (+30 types of labelled fields as possible during training also improves effectiveness, even if these fields are not available in the evaluation data (+13.5 citation parsing models. We further suggest that in future evaluations of reference parsers both evaluation data similar and dissimilar to the training data should be used for more meaningful evaluations.

READ FULL TEXT
research
05/12/2018

Citation Data-set for Machine Learning Citation Styles and Entity Extraction from Citation Strings

Citation parsing is fundamental for search engines within academia and t...
research
06/11/2019

EXmatcher: Combining Features Based on Reference Strings and Segments to Enhance Citation Matching

Citation matching is a challenging task due to different problems such a...
research
06/09/2020

Using BibTeX to Automatically Generate Labeled Data for Citation Field Extraction

Accurate parsing of citation reference strings is crucial to automatical...
research
05/23/2017

Reference String Extraction Using Line-Based Conditional Random Fields

The extraction of individual reference strings from the reference sectio...
research
02/04/2018

Machine Learning vs. Rules and Out-of-the-Box vs. Retrained: An Evaluation of Open-Source Bibliographic Reference and Citation Parsers

Bibliographic reference parsing refers to extracting machine-readable me...
research
02/04/2018

Evaluation and Comparison of Open Source Bibliographic Reference Parsers: A Business Use Case

Bibliographic reference parsing refers to extracting machine-readable me...

Please sign up or login with your details

Forgot password? Click here to reset