The Dawn of Today's Popular Domains: A Study of the Archived German Web over 18 Years

02/03/2017
by   Helge Holzmann, et al.
0

The Web has been around and maturing for 25 years. The popular websites of today have undergone vast changes during this period, with a few being there almost since the beginning and many new ones becoming popular over the years. This makes it worthwhile to take a look at how these sites have evolved and what they might tell us about the future of the Web. We therefore embarked on a longitudinal study spanning almost the whole period of the Web, based on data collected by the Internet Archive starting in 1996, to retrospectively analyze how the popular Web as of now has evolved over the past 18 years. For our study we focused on the German Web, specifically on the top 100 most popular websites in 17 categories. This paper presents a selection of the most interesting findings in terms of volume, size as well as age of the Web. While related work in the field of Web Dynamics has mainly focused on change rates and analyzed datasets spanning less than a year, we looked at the evolution of websites over 18 years. We found that around 70 are younger than a year, with an observed exponential growth in age as well as in size up to now. If this growth rate continues, the number of pages from the popular domains will almost double in the next two years. In addition, we give insights into our data set, provided by the Internet Archive, which hosts the largest and most complete Web archive as of today.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/16/2022

"Way back then": A Data-driven View of 25+ years of Web Evolution

Since the inception of the first web page three decades back, the Web ha...
research
06/09/2023

Mind2Web: Towards a Generalist Agent for the Web

We introduce Mind2Web, the first dataset for developing and evaluating g...
research
10/28/2021

A First Look at the Consolidation of DNS and Web Hosting Providers

Although the Internet continues to grow, it increasingly depends on a sm...
research
02/25/2019

Bootstrapping Domain-Specific Content Discovery on the Web

The ability to continuously discover domain-specific content from the We...
research
04/26/2019

Characterizing web pornography consumption from passive measurements

Web pornography represents a large fraction of the Internet traffic, wit...
research
07/22/2014

Artificial Life and the Web: WebAL Comes of Age

A brief survey is presented of the first 18 years of web-based Artificia...
research
05/31/2023

Beyond Rankings: Exploring the Impact of SERP Features on Organic Click-through Rates

Search Engine Result Pages (SERPs) serve as the digital gateways to the ...

Please sign up or login with your details

Forgot password? Click here to reset