CoVA: Context-aware Visual Attention for Webpage Information Extraction

10/24/2021
by   Anurendra Kumar, et al.
0

Webpage information extraction (WIE) is an important step to create knowledge bases. For this, classical WIE methods leverage the Document Object Model (DOM) tree of a website. However, use of the DOM tree poses significant challenges as context and appearance are encoded in an abstract manner. To address this challenge we propose to reformulate WIE as a context-aware Webpage Object Detection task. Specifically, we develop a Context-aware Visual Attention-based (CoVA) detection pipeline which combines appearance features with syntactical structure from the DOM tree. To study the approach we collect a new large-scale dataset of e-commerce websites for which we manually annotate every web element with four labels: product price, product title, product image and background. On this dataset we show that the proposed CoVA approach is a new challenging baseline which improves upon prior state-of-the-art methods.

READ FULL TEXT

page 2

page 8

research
05/07/2023

Context-Aware Chart Element Detection

As a prerequisite of chart data extraction, the accurate detection of ch...
research
11/08/2021

Learning Context-Aware Representations of Subtrees

This thesis tackles the problem of learning efficient representations of...
research
12/04/2020

Global Context Aware RCNN for Object Detection

RoIPool/RoIAlign is an indispensable process for the typical two-stage o...
research
01/07/2021

Simplified DOM Trees for Transferable Attribute Extraction from the Web

There has been a steady need to precisely extract structured knowledge f...
research
05/07/2019

Context-Aware Automatic Occlusion Removal

Occlusion removal is an interesting application of image enhancement, fo...
research
06/02/2021

Chunk Content is not Enough: Chunk-Context Aware Resemblance Detection for Deduplication Delta Compression

With the growing popularity of cloud storage, removing duplicated data a...
research
08/26/2023

How Can Context Help? Exploring Joint Retrieval of Passage and Personalized Context

The integration of external personalized context information into docume...

Please sign up or login with your details

Forgot password? Click here to reset