DeepAI AI Chat
Log In Sign Up

MementoMap Framework for Flexible and Adaptive Web Archive Profiling

by   Sawood Alam, et al.

In this work we propose MementoMap, a flexible and adaptive framework to efficiently summarize holdings of a web archive. We described a simple, yet extensible, file format suitable for MementoMap. We used the complete index of the comprising 5B mementos (archived web pages/files) to understand the nature and shape of its holdings. We generated MementoMaps with varying amount of detail from its HTML pages that have an HTTP status code of 200 OK. Additionally, we designed a single-pass, memory-efficient, and parallelization-friendly algorithm to compact a large MementoMap into a small one and an in-file binary search method for efficient lookup. We analyzed more than three years of MemGator (a Memento aggregator) logs to understand the response behavior of 14 public web archives. We evaluated MementoMaps by measuring their Accuracy using 3.3M unique URIs from MemGator logs. We found that a MementoMap of less than 1.5 comprehensive listing of all the unique original URIs) can correctly identify the presence or absence of 60 while maintaining 100


page 6

page 9


Where Did the Web Archive Go?

To perform a longitudinal investigation of web archives and detecting va...

Collecting 16K archived web pages from 17 public web archives

We document the creation of a data set of 16,627 archived web pages, or ...

Making Recommendations from Web Archives for "Lost" Web Pages

When a user requests a web page from a web archive, the user will typica...

Robots Still Outnumber Humans in Web Archives, But Less Than Before

To identify robots and humans and analyze their respective access patter...

Fully Automated HTML and Javascript Rewriting for Constructing a Self-healing Web Proxy

Over the last few years, the complexity of web applications has increase...

Archive Assisted Archival Fixity Verification Framework

The number of public and private web archives has increased, and we impl...

FastWARC: Optimizing Large-Scale Web Archive Analytics

Web search and other large-scale web data analytics rely on processing a...