Extraction of Relevant Images for Boilerplate Removal in Web Browsers

12/17/2019
by   Joy Bose, et al.
0

Boilerplate refers to unwanted and repeated parts of a webpage (such as ads or table of contents) that distracts the user from reading the core content of the webpage, such as a news article. Accurate detection and removal of boilerplate content from a webpage can enable the users to have a clutter free view of the webpage or news article. This can be useful in features like reader mode in web browsers. Current implementations of reader mode in web browsers such as Firefox, Chrome and Edge perform reasonably well for textual content in webpages. However, they are mostly heuristic based and not flexible when the webpage content is dynamic. Also they often do not perform well for removing boilerplate content in the form of images and multimedia in webpages. For detection of boilerplate images, one needs to have knowledge of the actual layout of the images in the webpage, which is only possible when the webpage is rendered. In this paper we discuss some of the issues in relevant image extraction. We also present the design of a testing framework to measure accuracy and a classifier to extract relevant images by leveraging a headless browser solution that gives the rendering information for images.

READ FULL TEXT
research
11/08/2019

Semi-Supervised Method using Gaussian Random Fields for Boilerplate Removal in Web Browsers

Boilerplate removal refers to the problem of removing noisy content from...
research
04/10/2021

A Web Infrastructure for Certifying Multimedia News Content for Fake News Defense

In dealing with altered visual multimedia content, also referred to as f...
research
08/31/2021

Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection

User-generated content (e.g., tweets and profile descriptions) and share...
research
12/27/2018

Chart-Text: A Fully Automated Chart Image Descriptor

Images greatly help in understanding, interpreting and visualizing data....
research
03/06/2023

Implementation of a noisy hyperlink removal system: A semantic and relatedness approach

As the volume of data on the web grows, the web structure graph, which i...
research
06/21/2022

Broken News: Making Newspapers Accessible to Print-Impaired

Accessing daily news content still remains a big challenge for people wi...
research
10/03/2012

Logical segmentation for article extraction in digitized old newspapers

Newspapers are documents made of news item and informative articles. The...

Please sign up or login with your details

Forgot password? Click here to reset