A principled methodology for comparing relatedness measures for clustering publications

01/21/2019
by   Ludo Waltman, et al.
0

There are many different relatedness measures, based for instance on citation relations or textual similarity, that can be used to cluster scientific publications. We propose a principled methodology for evaluating the accuracy of clustering solutions obtained using these relatedness measures. We formally show that the proposed methodology has an important consistency property. The empirical analyses that we present are based on publications in the fields of cell biology, condensed matter physics, and economics. Using the BM25 text-based relatedness measure as evaluation criterion, we find that bibliographic coupling relations yield more accurate clustering solutions than direct citation relations and co-citation relations. The so-called extended direct citation approach performs similarly to or slightly better than bibliographic coupling in terms of the accuracy of the resulting clustering solutions. The other way around, using a citation-based relatedness measure as evaluation criterion, BM25 turns out to yield more accurate clustering solutions than other text-based relatedness measures.

READ FULL TEXT
research
02/11/2017

Citation-based clustering of publications using CitNetExplorer and VOSviewer

Clustering scientific publications in an important problem in bibliometr...
research
04/10/2020

Return to basics: Clustering of scientific literature using structural information

Scholars frequently employ relatedness measures to estimate the similari...
research
10/29/2021

Generalization of bibliographic coupling and co-citation using the node split network

Bibliographic coupling (BC) and co-citation (CC) are the two most common...
research
10/30/2018

The dispersion of the citation distribution of top scientists' publications

This work explores the distribution of citations for the publications of...
research
05/17/2021

A Measure of Research Taste

Researchers are often evaluated by citation-based metrics. Such metrics ...
research
05/10/2020

Frequently Co-cited Publications: Features and Kinetics

Co-citation measurements can reveal the extent to which a concept repres...
research
10/13/2018

Measuring Swampiness: Quantifying Chaos in Large Heterogeneous Data Repositories

As scientific data repositories and filesystems grow in size and complex...

Please sign up or login with your details

Forgot password? Click here to reset