Full-Text and URL Search Over Web Archives

08/03/2021
by   Miguel Costa, et al.
0

Web archives are a historically valuable source of information. In some respects, web archives are the only record of the evolution of human society in the last two decades. They preserve a mix of personal and collective memories, the importance of which tends to grow as they age. However, the value of web archives depends on their users being able to search and access the information they require in efficient and effective ways. Without the possibility of exploring and exploiting the archived contents, web archives are useless. Web archive access functionalities range from basic browsing to advanced search and analytical services, accessed through user-friendly interfaces. Full-text and URL search have become the predominant and preferred forms of information discovery in web archives, fulfilling user needs and supporting search APIs that feed complex applications. Both full-text and URL search are based on the technology developed for modern web search engines, since the Web is the main resource targeted by both systems. However, while web search engines enable searching over the most recent web snapshot, web archives enable searching over multiple snapshots from the past. This means that web archives have to deal with a temporal dimension that is the cause of new challenges and opportunities, discussed throughout this chapter.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/24/2019

From Search Engines to Search Services: An End-User Driven Approach

The World Wide Web is a vast and continuously changing source of informa...
research
01/28/2017

How to Search the Internet Archive Without Indexing It

Significant parts of cultural heritage are produced on the web during th...
research
07/09/2020

On the Social and Technical Challenges of Web Search Autosuggestion Moderation

Past research shows that users benefit from systems that support them in...
research
10/21/2020

Literature Review of Computer Tools for the Visually Impaired: a focus on Search Engines

A sudden reliance on the internet has resulted in the global standardiza...
research
05/30/2018

DATA:SEARCH'18 – Searching Data on the Web

This half day workshop explores challenges in data search, with a partic...
research
12/09/2021

Feature Modulation to Improve Struggle Detection in Web Search: A Psychological Approach

Searcher struggle is important feedback to Web search engines. Existing ...
research
05/16/2015

Cognitive Development of the Web

The sociotechnological system is a system constituted of human individua...

Please sign up or login with your details

Forgot password? Click here to reset