Provenance for Linguistic Corpora Through Nanopublications

06/11/2020
by   Timo Lek, et al.
0

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without significant manual work. Moreover, these annotations might have adapted and might have evolved into different versions, making it challenging for researchers to know the data's provenance and merge it with other annotated corpora. In other words, these variations affect the interoperability between existing corpora. This paper addresses this issue with a case study on event annotated corpora and by creating a new, more interoperable representation of this data in the form of nanopublications. We demonstrate how linguistic annotations from separate corpora can be merged through a similar format to thereby make annotation content simultaneously accessible. The process for developing the nanopublications is described, and SPARQL queries are performed to extract interesting content from the new representations. The queries show that information of multiple corpora can now be retrieved more easily and effectively with the automated interoperability of the information of different corpora in a uniform data format.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/25/2022

Graph Querying for Semantic Annotations

This paper presents how the online tool GREW-MATCH can be used to make q...
research
01/19/2022

ASL Video Corpora Sign Bank: Resources Available through the American Sign Language Linguistic Research Project (ASLLRP)

The American Sign Language Linguistic Research Project (ASLLRP) provides...
research
11/22/2020

Standardizing linguistic data: method and tools for annotating (pre-orthographic) French

With the development of big corpora of various periods, it becomes cruci...
research
08/21/2018

Analysis of Speeches in Indian Parliamentary Debates

With the increasing usage of the internet, more and more data is being d...
research
04/26/2019

Producing Corpora of Medieval and Premodern Occitan

At a time when the quantity of - more or less freely - available data is...
research
07/01/2019

PAGAN: Video Affect Annotation Made Easy

How could we gather affect annotations in a rapid, unobtrusive, and acce...
research
03/03/2020

Seshat: A tool for managing and verifying annotation campaigns of audio data

We introduce Seshat, a new, simple and open-source software to efficient...

Please sign up or login with your details

Forgot password? Click here to reset