Distributed Holistic Clustering on Linked Data

08/30/2017
by   Markus Nentwig, et al.
0

Link discovery is an active field of research to support data integration in the Web of Data. Due to the huge size and number of available data sources, efficient and effective link discovery is a very challenging task. Common pairwise link discovery approaches do not scale to many sources with very large entity sets. We here propose a distributed holistic approach to link many data sources based on a clustering of entities that represent the same real-world object. Our clustering approach provides a compact and fused representation of entities, and can identify errors in existing links as well as many new links. We support a distributed execution of the clustering approach to achieve faster execution times and scalability for large real-world data sets. We provide a novel gold standard for multi-source clustering, and evaluate our methods with respect to effectiveness and efficiency for large data sets from the geographic and music domains.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/22/2019

Integrating Information About Entities Progressively

Users often have to integrate information about entities from multiple d...
research
04/28/2016

Exploiting Source-Object Network to Resolve Object Conflicts in Linked Data

Considerable effort has been made to increase the scale of Linked Data. ...
research
03/03/2018

MaskLink: Efficient Link Discovery for Spatial Relations via Masking Areas

In this paper, we study the problem of spatial link discovery (LD), focu...
research
03/15/2021

Online Topic-Aware Entity Resolution Over Incomplete Data Streams (Technical Report)

In many real applications such as the data integration, social network a...
research
07/06/2023

JSONoid: Monoid-based Enrichment for Configurable and Scalable Data-Driven Schema Discovery

Schema discovery is an important aspect to working with data in formats ...
research
06/05/2019

VoIDext: Vocabulary and patterns for enhancing interoperable datasets with virtual links

Semantic heterogeneity remains a problem when interoperating with data f...
research
03/07/2016

TruthDiscover: Resolving Object Conflicts on Massive Linked Data

Considerable effort has been made to increase the scale of Linked Data. ...

Please sign up or login with your details

Forgot password? Click here to reset