Wikipedia Text Reuse: Within and Without

12/21/2018
by   Milad Alshomary, et al.
0

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy and paste, we employ state-of-the-art text reuse detection technology, scaling it for the first time to process the entire Wikipedia as part of a distributed retrieval pipeline. We further report on a pilot analysis of the 100 million reuse cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered. Text reuse inside Wikipedia gives rise to new tasks such as article template induction, fixing quality flaws due to inconsistencies arising from asynchronous editing of reused passages, or complementing Wikipedia's ontology. Text reuse outside Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia's influence on the web. To foster future research into these tasks, and for reproducibility's sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely available.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/22/2023

TEIMMA: The First Content Reuse Annotator for Text, Images, and Math

This demo paper presents the first tool to annotate the reuse of text, i...
research
07/26/2023

Measuring Americanization: A Global Quantitative Study of Interest in American Topics on Wikipedia

We conducted a global comparative analysis of the coverage of American t...
research
12/22/2021

STEREO: Scientific Text Reuse in Open Access Publications

We present the Webis-STEREO-21 dataset, a massive collection of Scientif...
research
04/21/2020

Use of Wikipedia categories on information retrieval research: a brief review

Wikipedia categories, a classification scheme built for organizing and d...
research
05/10/2021

Wiki-Reliability: A Large Scale Dataset for Content Reliability on Wikipedia

Wikipedia is the largest online encyclopedia, used by algorithms and web...
research
02/08/2023

Reception Reader: Exploring Text Reuse in Early Modern British Publications

The Reception Reader is a web tool for studying text reuse in the Early ...
research
12/18/2021

The Web Is Your Oyster – Knowledge-Intensive NLP against a Very Large Web Corpus

In order to address the increasing demands of real-world applications, t...

Please sign up or login with your details

Forgot password? Click here to reset