Merging error analysis of name disambiguation based on author similarity

11/04/2017
by   Zheng Xie, et al.
0

Falsely identifying different authors as one is called merging error in the name disambiguation of coauthorship networks. Research on the measurement and distribution of merging errors helps to collect high quality coauthorship networks. In the aspect of measurement, we provide a Bayesian model to measure the errors through author similarity. We illustratively use the model and coauthor similarity to measure the errors caused by initial-based name disambiguation methods. The empirical result on large-scale coauthorship networks shows that using coauthor similarity cannot increase the accuracy of disambiguation through surname and the initial of the first given name. In the aspect of distribution, expressing coauthorship data as hypergraphs and supposing the merging error rate is proper to hyperdegree with an exponent, we find that hypergraphs with a range of network properties highly similar to those of low merging error hypergraphs can be constructed from high merging error hypergraphs. It implies that focusing on the error correction of high hyperdegree nodes is a labor- and time-saving approach of improving the data quality for coauthorship network analysis.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/13/2020

A class of ie-merging functions

We describe a general class of ie-merging functions and pose the problem...
research
11/02/2014

High Dynamic Range Imaging by Perceptual Logarithmic Exposure Merging

In this paper we emphasize a similarity between the Logarithmic-Type Ima...
research
12/10/2020

Descriptive and Predictive Analysis of Aggregating Functions in Serverless Clouds: the Case of Video Streaming

Serverless clouds allocate multiple tasks (e.g., micro-services) from mu...
research
07/11/2012

Similarity-Driven Cluster Merging Method for Unsupervised Fuzzy Clustering

In this paper, a similarity-driven cluster merging method is proposed fo...
research
05/24/2020

The effect of measurement error on clustering algorithms

Clustering consists of a popular set of techniques used to separate data...
research
07/30/2014

Merging and Shifting of Images with Prominence Coefficient for Predictive Analysis using Combined Image

Shifting of objects in an image and merging many images after appropriat...
research
09/26/2018

Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection

Grammatical error correction, like other machine learning tasks, greatly...

Please sign up or login with your details

Forgot password? Click here to reset